Pith. sign in

Paper Citation Record · LEDGER

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition

As of 6 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 0 inbound Pith citation observations for arXiv:2604.23173.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.23173 v1

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T08:50:27.871886Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

84 of 84 outbound references displayed

  • verified exact11
  • verified fuzzy73
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5fd60d4-5da1-42ae-b3e7-634e9d122bfd · outbound

This paper cites https : / / github.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition https : / / github

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.112354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:a71159e2604b622849ed653ae14d45d70d1a30a80e74765471079f6373c31e94

Observation e9b4570f-374a-49b7-92d7-b5057b6746cb · outbound

This paper cites One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.040721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:d9e05c640649b3d4f446b311cc95213c24c124146f464f9eb6176fca0c9211bc

Observation 1c8cdd0e-be70-4ef3-a57f-28475efef9e0 · outbound

This paper cites Frozen in Time: A Joint Video and Image Encoder for End- to-End Retrieval.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Frozen in Time: A Joint Video and Image Encoder for End- to-End Retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.089947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:bdbb323cf76e97bea6b160623c4b3b436f1c9a758cf6623399bf7360acaf70e7

Observation f26f2c02-c9c7-4403-8da3-743bb2554930 · outbound

This paper cites XMem++: Production-Level Video Segmentation from Few Annotated Frames.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition XMem++: Production-Level Video Segmentation from Few Annotated Frames

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.095209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:c479c6289e40f2860cf96bcf03ef44862db114650ac3cb1a566c0fdefd3c383b

Observation 6759dafd-8868-4720-a536-791728ff7481 · outbound

This paper cites Corefer- ence Resolution through a Seq2Seq Transition-Based System.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Corefer- ence Resolution through a Seq2Seq Transition-Based System

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.043072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:fbb7aca061de4491f76f6adc3c201decb4feefc9e00f300eb3824b48536b45c0

Observation 01b4c86a-3986-4123-8c92-63e253b13e1c · outbound

This paper cites Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.052519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:94623014f6fbe5709b514e3af9325962957038e2d8838e85a87fba24405ddc42

Observation c71bc4cd-8fde-46c5-a55c-0526d8728d55 · outbound

This paper cites Joint Multimedia Event Extraction from Video and Article.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Joint Multimedia Event Extraction from Video and Article

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.228888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:d25a48b1e2261cdcb7bd87cfb08642f57b92b37beca30ee8f8e820e19f673540

Observation 7eebcc19-a9c6-4642-b3ee-d7820a122961 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:52:36.167809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:d9b0482b091df36516d908fd70347d53e2d1fb278ebdb5ae3a170bd87b6faaed

Observation 767e1352-4543-45af-b6e0-64216e8d3dde · outbound

This paper cites ShareGPT4Video: Improving Video Understand- ing and Generation with Better Captions.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition ShareGPT4Video: Improving Video Understand- ing and Generation with Better Captions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.144911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:e1108fa6c3af64e997ad8feef6f3403ebd948f81feee1168516c9666a46531bf

Observation 38d54ce3-335a-452d-9fb3-7ef616a232c9 · outbound

This paper cites V AST: A Vision-Audio- Subtitle-Text Omni-Modality Foundation Model and Dataset.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition V AST: A Vision-Audio- Subtitle-Text Omni-Modality Foundation Model and Dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.177079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:93cd71c65036684a72bd6b1e0db8fa1541e081796a6a0b80da85e68f1c446b7c

Observation f02f9336-0ece-4db3-967b-0c49fc8fd30b · outbound

This paper cites XMem: Long- Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition XMem: Long- Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.191903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:cef97ca4cc43e35c116d4709b2ea4e4eeec7fd8e367126a6e1eb1ab0a5a6e6be

Observation 119424ae-d507-4a02-a73c-9eb699106c1d · outbound

This paper cites Segment and Track Anything.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Segment and Track Anything

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:11.807561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:23351b8cd188249db8e85268b00e9ad6fe837d3ba2445844f34b9190898d3e41

Observation 3fae2782-2031-47a0-835c-4875430e5a47 · outbound

This paper cites Zero-Shot Video Question Answering with Procedural Programs.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Zero-Shot Video Question Answering with Procedural Programs

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.114933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:31b47ff5abb2ecfc870fc41fd4f8aab19822fb2c9223fa33b9153a35eb16bdf1

Observation ebe15fdc-e2b7-4bc7-8045-685d8c28d27f · outbound

This paper cites Word-Level Coreference Resolu- tion.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Word-Level Coreference Resolu- tion

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.111336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:520ac8cb9e1a80752969afe2bc3011ec4c985d31b59a9ba05f06b795c1e446ee

Observation 214f6afe-4e1a-42f3-a79b-c9e0099d7b46 · outbound

This paper cites DAPS: Deep Action Proposals for Action Understanding.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition DAPS: Deep Action Proposals for Action Understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.120838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:9e69e526912ac11b09990e177d8550ab12e534f74b4dfaf83f61839261a99d1d

Observation 67dc4c8f-9d1e-4c53-8e4b-20ee3b7e09ca · outbound

This paper cites SlowFast Networks for Video Recognition.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition SlowFast Networks for Video Recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.189019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:b466784784a0e4c5ac735ace6d1e88201b83e09d1c2d848049694787f76c5da6

Observation dee19dc4-5c2e-46f4-bae7-be3e3a394e14 · outbound

This paper cites Video Action Transformer Network.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Video Action Transformer Network

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.210833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:860e82f5782475f6e68399792ac2a3226c158d6db28b8cbab250d9dfd2ae8415

Observation c27d6c7b-6794-4172-b272-4b384c50cce8 · outbound

This paper cites Semi-Supervised Multimodal Coreference Resolution in Im- age Narrations.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Semi-Supervised Multimodal Coreference Resolution in Im- age Narrations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.072751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:9506f17cef5846b7edb1f524c9ab2191c404cdb272e65ffbe2ab0ed1d14e884a

Observation 83027b6e-c2ac-4f63-929d-038cb2a6290d · outbound

This paper cites Who Are You Referring To? Coreference Resolution in Image Narrations.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Who Are You Referring To? Coreference Resolution in Image Narrations

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.194936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:9367276ec976b82b03c316daa4c40e3cb60838b802954f9c618608308d87a2e4

Observation a26fbe97-bf40-4db1-a69d-2b6dddd5469d · outbound

This paper cites AGQA: A Benchmark for Compositional Spatio-Temporal Reasoning.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition AGQA: A Benchmark for Compositional Spatio-Temporal Reasoning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.097846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:323978792ed46660904f95ab124cbf48c2e6898e4c72ba8c61a32e9a06104e1e

Observation b3fdf8a4-28ef-4508-b572-cfee712e33bd · outbound

This paper cites Fast Temporal Activity Proposals for Efficient De- tection of Human Actions in Untrimmed Videos.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Fast Temporal Activity Proposals for Efficient De- tection of Human Actions in Untrimmed Videos

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.100431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:46aa1ea99e5ce27c7f8f2e1d964e0ef91e3d92a2d8970c7a023f531f742c87a8

Observation 8a2ab035-2da3-4b5c-a2ea-b7e55174ded1 · outbound

This paper cites VTIMELLM: Empower LLM to Grasp Video Mo- mentsVideo Action Transformer Network.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition VTIMELLM: Empower LLM to Grasp Video Mo- mentsVideo Action Transformer Network

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.103011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:ae4ddf6a5e2ad81dc73a017a4563b893a03ec5741c17f35cd9cf27a7344978dc

Observation b806f9ed-f97a-4ab1-a9ea-5a112b30d0b0 · outbound

This paper cites Action Genome: Actions as Compositions of Spatio- Temporal Scene Graphs.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Action Genome: Actions as Compositions of Spatio- Temporal Scene Graphs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.153002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:fec6a7796d6fd80a0b1af3eecf97d0ee3747cfa4eda35cf5ac9ce48e236c08cd

Observation 29f8e0d1-3f5a-4d90-9289-3cb86b4cab97 · outbound

This paper cites ReferItGame: Referring to Objects in Pho- tographs of Natural Scenes.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition ReferItGame: Referring to Objects in Pho- tographs of Natural Scenes

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.116560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:6c07a2675464d8575f1112d24a9a3d1a0070c94e5b541399c17e5bdd54ed90ac

Observation f6ef0d16-4bc6-46f6-b17e-91e79511920e · outbound

This paper cites Grounded Video Situation Recognition.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Grounded Video Situation Recognition

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.174225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:bcc2f344e141d2d977b351b75446a68cfb6256254c991ea84aada263f1ec7fe6

Observation 01e556e0-a124-4720-a4dc-5386ad82038b · outbound

This paper cites Kingma and Jimmy Ba.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Kingma and Jimmy Ba

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.218839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:7064f8245d8d9a77a166f10b76a3a6aeae4beb0087321422a5f5d0d01ed35338

Observation acc491be-d0f4-4bdb-ae1c-1d5988e66a5a · outbound

This paper cites Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations.International Journal of Computer Vision (IJCV), 123(1):32–73.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations.International Journal of Computer Vision (IJCV), 123(1):32–73

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.223783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:d0d12476f36cc7fed621dafd5221edaafd8d23d93399a46872f10bd302d61149

Observation 8d6ccf2e-2812-4b8d-8ae5-0f470269bbe9 · outbound

This paper cites SRTube: Video-Language Pre-Training with Action-Centric Video Tube Features and Semantic Role Labeling.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition SRTube: Video-Language Pre-Training with Action-Centric Video Tube Features and Semantic Role Labeling

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.213746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:e4610c72dfb1e6284cd111b9306c2bb3fe010410b983d22fca571a185d78f825

Observation 43975666-5855-4df1-a52a-50d1b6998428 · outbound

This paper cites TVQA: Localized, Compositional Video Question Answering.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition TVQA: Localized, Compositional Video Question Answering

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.208136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:03fa1c213e543c8809c2394cfb73cb5aa50196f096bc852a8508493ad1229e92

Observation b265e963-c916-49a7-bf23-ee5953bf9386 · outbound

This paper cites BLIP- 2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition BLIP- 2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.087310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:1d0e1371dc0f3142b4075167cc2fb03fe2bb108d9552662d0cb1ff863f258fe9

Observation 928ae01a-504a-4a23-b085-ae0cc28929cb · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition VILA: On Pre-training for Visual Language Models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.216402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:87b0177e8d40f3c9e574b9f75b56ec2b8665ff1a8a6fe0304840cf0e77b1ce9d

Observation 66527e88-d5a1-42a0-8c01-697c15cf6ac1 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.150019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:390a4a22834105ebf13a168685a08704e075425e50394cba7834836f77865406

Observation 67b043ed-f8c5-443c-b858-eb5a247b522e · outbound

This paper cites HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.046283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:a8da170c29cf946d673e2dda6a94f5f073ad800a811c1b77a620a655d3015ef5

Observation 69887a59-b9d2-4306-98ad-b43f2acfe038 · outbound

This paper cites MoMA: Multi-Object Multi-Actor Activity Pars- ing.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition MoMA: Multi-Object Multi-Actor Activity Pars- ing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.041164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:03bcdcac0f87123c579f7a2a7b445ac10b1926b2efc3e3d017b71b88f2f0c000

Observation 7c01aaa3-64fa-4af2-934a-7604b493ac9e · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.055258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:02e0e0536e5ee6eeff4f935df2d3a8021b1bf26f8449adc63256bac95b62f2f8

Observation c2927b7c-2082-4279-859d-a7d2ccf2b0a4 · outbound

This paper cites Major Entity Identification: A Generalizable Alternative to Coreference Resolution.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Major Entity Identification: A Generalizable Alternative to Coreference Resolution

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.144347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:583dc938b7478fbc5bc08eb278399d866dc0481511e43edb6407d04e581bc3b9

Observation a5fe1333-ec10-42ea-8935-3cdc449df7de · outbound

This paper cites Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:11.798761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:f090561de3dbede09c9b18e9eac02ed77ef6e1a50ec1434b52674c22b111c6d7

Observation 92fa1bf3-ae89-4bbe-b20a-e9b5c90b7aaf · outbound

This paper cites HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.104511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:8183c2156eaf3a0748c5aa05288dad3b49243c28f9279f2e0f3f9ecf65cc3cb0

Observation 4b3e6cc8-b14e-4638-8b36-cfec3d5b77ee · outbound

This paper cites Which Coref- erence Evaluation Metric Do You Trust? A Proposal for a Link-Based Entity Aware Metric.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Which Coref- erence Evaluation Metric Do You Trust? A Proposal for a Link-Based Entity Aware Metric

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.064126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:11f3302fbafacae2ec255ebff51727a27c55fdc043420b32c398fe095bec90f6

Observation 0ed80c0f-0db6-4f7b-9034-77b733219f79 · outbound

This paper cites VideoGLAMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition VideoGLAMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.200371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:ab4ac3353d12a41b5b64f100688f1a550fcbea032b64a28e73f1beb91e6fcbc6

Observation 36ca9cc9-1f11-4485-b37d-b63476214d6a · outbound

This paper cites PG-Video-LLaVA: Pixel Grounding Large Video-Language Models.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition PG-Video-LLaVA: Pixel Grounding Large Video-Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:11.774123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:4f6ad91af2b1f36f181204e2154ca3f3bfebbe130f4ef33f9c758d8d4bb01f6a

Observation f8c1142b-96fd-47f9-b722-c0084bee742b · outbound

This paper cites HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video Understanding.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video Understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.221537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:db3f005aefb3ac6b7cfd640162f71819f53a4fcb4592f5add9420701f8f9fea9

Observation 2dc87950-735b-4c4f-bdaa-b080e31da4e5 · outbound

This paper cites Identity- Aware Multi-Sentence Video Description.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Identity- Aware Multi-Sentence Video Description

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.162427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:914e1486c400a39d6448a86ba84b53dc289617057a8e70f843bb734f01f21240

Observation 2b1db321-013e-4e36-9cdc-b0edfb449562 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.096433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:a384eb44fa1d01d8e2911ad8069099acd4fd6256757e434538dd234392a159a8

Observation 0bc7ec90-94c3-40f7-9da5-038858828236 · outbound

This paper cites Qwen2.5-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution.Blog post: https://qwenlm.github.io/blog/qwen2.5-vl/.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Qwen2.5-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution.Blog post: https://qwenlm.github.io/blog/qwen2.5-vl/

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.156059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:f886a76516dd6bff89b8a677453ea87fa8268b84491b54d0dcce3f06eacd7a46

Observation e2e9ac69-14a1-427a-931c-a120ccb61b2d · outbound

This paper cites MICAP: A Unified Model for Identity- Aware Movie Descriptions.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition MICAP: A Unified Model for Identity- Aware Movie Descriptions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.159750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:8fa8ebb5951a9985c51004df00e7874554c0966f45d60c629a66a7271418b211

Observation e8da2b91-4671-4350-9274-29708ceac499 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition SAM 2: Segment Anything in Images and Videos

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:31:11.818680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:237484515f4079c041b17e4507de0105245640aadd06b26838b9d103a8cf85b3

Observation a0f4996a-2bc4-4385-bc85-a3f65f7ea9c0 · outbound

This paper cites Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.165410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:c915617b28a19a703f68c7bc1cd06bc0cf2c8cc873b359af69f747aa2b261510

Observation f8d5395d-b575-43d2-a51a-a43a152573ed · outbound

This paper cites Movie Description.International Journal of Computer Vision (IJCV), 123:94–120.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Movie Description.International Journal of Computer Vision (IJCV), 123:94–120

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.183419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:d1c5e49a604fec9738ef14f69b04b0c6e82e75c9164b77706dbdfdcf63bf2255

Observation c19619df-c08c-447f-a8a3-ff06349cf4ee · outbound

This paper cites Video Object Grounding using Semantic Roles in Language Description.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Video Object Grounding using Semantic Roles in Language Description

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.106875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:3e0f5cb1baad751b2ca21a792d790051f384515358af97d0313651cdf9afdec3

Observation b2ca91a1-1a68-4777-8b50-07714d9f30cc · outbound

This paper cites Visual Semantic Role Labeling for Video Understanding.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Visual Semantic Role Labeling for Video Understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.109259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:415f5c937f058f8b7f2c9f1aeee62633abd8bc157056f50f2e913502eb6d96d3

Observation 5846eb1d-9e95-412a-bfd5-a1708b1008ab · outbound

This paper cites VELOC- ITI: Can Video-Language Models Bind Semantic Concepts through Time? InConference on Computer Vision and Pat- tern Recognition (CVPR).

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition VELOC- ITI: Can Video-Language Models Bind Semantic Concepts through Time? InConference on Computer Vision and Pat- tern Recognition (CVPR)

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.139359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:00db019765b8f90e0a1a63af29c636e1412bcedb09c8ea1464db39e0b46d61f4

Observation 3ea43469-2cba-4504-9be5-d34441bf71e1 · outbound

This paper cites Effi- cient Parameter-Free Clustering Using First Neighbor Rela- tions.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Effi- cient Parameter-Free Clustering Using First Neighbor Rela- tions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.231552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:18ec112da6f977f0eaad231dbc34b61edc50ad2008c3b8432c0c3abc6d180f8b

Observation e06e8275-7c63-4537-a501-26f0eb69d98c · outbound

This paper cites End-to-End Generative Pretraining for Multimodal Video Captioning.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition End-to-End Generative Pretraining for Multimodal Video Captioning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.136578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:d23bee9fa92a6839ae9f18d3fa143fbf833ce783a967913645f0db773904e31a

Observation 196b3e67-4de3-459c-adf2-537168ee7409 · outbound

This paper cites Temporal Action Localization in Untrimmed Videos via Multi-Stage CNNs.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Temporal Action Localization in Untrimmed Videos via Multi-Stage CNNs

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.142355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:3e21dc404962768238643b5c7bdea39fbc2deb8214e986a329587f31882b9dd2

Observation 2d72c2cb-3914-404a-bdbb-35176f7b7336 · outbound

This paper cites Actor-Centric Relation Network.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Actor-Centric Relation Network

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.147302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:ab04c14e2967c6db194902a2812d2f014aa8151a1030f47c2d7fba59c8d37350

Observation cb637cd7-02cf-4ebb-ab84-343549e88ac6 · outbound

This paper cites Unbiased Scene Graph Generation from Biased Training.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Unbiased Scene Graph Generation from Biased Training

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.171773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:3c44163dec21a2fde71497bd4cf81e7c396698aecc035f2667808180a5cd52a8

Observation 7671a9fc-1460-457d-815a-24df8c021001 · outbound

This paper cites Long Term Spatio-Temporal Modeling for Action Detection.Com- puter Vision and Image Understanding (CVIU), 210.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Long Term Spatio-Temporal Modeling for Action Detection.Com- puter Vision and Image Understanding (CVIU), 210

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.180029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:6ded76c4b5ad203e4db94a9512c667403ceae2f50f2e48adfc73490f2ceceb5e

Observation 39124621-bd7d-405d-ac3f-12aa3f6ce112 · outbound

This paper cites MovieQA: Un- derstanding Stories in Movies through Question-Answering.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition MovieQA: Un- derstanding Stories in Movies through Question-Answering

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.133405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:43b1edaeddf7d9e519b574a9a15a13c2bdbdcc040906cfab92d44ca11173999b

Observation 1e6903fd-efce-4b99-a116-d85e7b869e96 · outbound

This paper cites On Generalization in Coref- erence Resolution.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition On Generalization in Coref- erence Resolution

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.130095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:43169f647b0355702d3ce9233eea6cd4e5b7763f1893103429b6100d437ca52b

Observation 1cd80ef2-6eb8-4843-baea-d03e5eeb964f · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:31:11.788353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:51b29170d472cb9d698010ecd9a4470d1f3c5a022e3166e8d405b8be61f5ff62

Observation 31aa72e1-dd44-4e54-9546-29c6c840f21d · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Lawrence Zitnick, and Devi Parikh

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.168728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:b667d4e6593c200867d44f6a4603377cbaf680703d0034b2a617f8910546ede3

Observation fe71fc0c-9156-40a5-b457-f4f455bc76f8 · outbound

This paper cites Effec- tively Leveraging CLIP for Generating Situational Summaries of Images and Videos.IJCV.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Effec- tively Leveraging CLIP for Generating Situational Summaries of Images and Videos.IJCV

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.035037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:4105b180ec4fd8d9c252ee880c430c269d06b8b9f9fdd07c0187c877e2a8da3f

Observation 30d3bcc8-0f56-4fc2-a711-ed08d1fcfa21 · outbound

This paper cites MovieGraphs: Towards Understanding Human-Centric Situations from Videos.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition MovieGraphs: Towards Understanding Human-Centric Situations from Videos

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.117798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:8e1286768fb676c9652645a18fc85d74bb694e32fd21f43395851c9f50717bb9

Observation e1fe5fc1-dec3-463a-bbc3-86db0d6b851e · outbound

This paper cites Yoloe: Real-time seeing anything.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Yoloe: Real-time seeing anything

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:11.863762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:795bd46f9c7a8384050c39c8983e1f48c5616d303ce085894efe517970c55c33

Observation 6d63972a-4b40-421f-b0e7-c2224d9c66b8 · outbound

This paper cites Temporal Segment Networks: Towards Good Practices for Deep Action Recog- nition.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Temporal Segment Networks: Towards Good Practices for Deep Action Recog- nition

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.099365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:a1b97c7975a422a5d0c116aa12d4858af54253267a2bad8829a1dd1405940f54

Observation a65125e1-11c6-4ed3-9600-e173705ebcfc · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:31:11.871197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:00d35a2fc3dd04d413371997b6737859a05932b148e7463e70a972afbecd0ce6

Observation 9241731f-8b22-4df2-9138-2d8297dd70bf · outbound

This paper cites Demonstration Meets Typed Events: Type Specific Video Semantic Role Labeling via Multimodal Prompting and Retrieval.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Demonstration Meets Typed Events: Type Specific Video Semantic Role Labeling via Multimodal Prompting and Retrieval

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.065507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:a2962aae9b930e075616ff505764d9bc4a7f6d7ad978f73a553f041cd94a6268

Observation 64dbb806-2a2a-4201-8e91-4382ba232fee · outbound

This paper cites Long-Term Feature Banks for Detailed Video Understanding.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Long-Term Feature Banks for Detailed Video Understanding

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.137739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:5ff8b5cf563f8973e39a3d9cf2e9011c13387654b3b2505b3e76cf3fa7bd6708

Observation a9ecbf17-778e-45ab-8246-33d12982c418 · outbound

This paper cites MSR-VTT: A Large Video Description Dataset for Bridging Video and Language.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition MSR-VTT: A Large Video Description Dataset for Bridging Video and Language

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.197356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:3d6e9c067820c9408d925c4ba44d59844fb8c271d70cd4c6f1af9bfb29b34921

Observation 7096acac-5063-44bf-9ecc-07f0b7ec918c · outbound

This paper cites TubeDETR: Spatio-Temporal Video Grounding with Transformers.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition TubeDETR: Spatio-Temporal Video Grounding with Transformers

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.226216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:b66c935058a36cfa4eb5c6ee4b9d753a2666181e099b4032edb5655279d2fd72

Observation 5bbd514c-21f6-41c5-b99f-4cea37362700 · outbound

This paper cites Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.127454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:14b1238dd2ad4c553c5619e29204c638b7ad6f3baccd3ac8c579069a42a89a1f

Observation cad60622-b58b-412d-a5c3-55d6378a9919 · outbound

This paper cites Video Event Extraction via Tracking Vi- sual States of Arguments.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Video Event Extraction via Tracking Vi- sual States of Arguments

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.075574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:ddd6305685aae1c953e68c4f213c437efda677af4c2c1643dac4defa78b99afa

Observation 9f126d71-9aa3-43ad-b798-6358d6c3dcea · outbound

This paper cites Track Anything: Segment Anything Meets Videos.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Track Anything: Segment Anything Meets Videos

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:31:11.879222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:63a07fea3f462fc2963f55098d1fb272aa6ea9a1b17e48e8fb4e868bc33c9ad9

Observation e40233af-9e7f-4d23-9105-e79b645cb135 · outbound

This paper cites Panoptic Video Scene Graph Generation.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Panoptic Video Scene Graph Generation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.202645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:e7913adc05b12b3468f5197a05ba787b03905ba3996ebd6d0e319861cc6b5ec6

Observation aaec3d26-26c8-4d62-a79b-42d8e42aa024 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:31:11.855018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:481ca94dce1aa54473a16b3c91fb9294ab1e6b103fd2388b3c0813fb380471d5

Observation 16b9b879-28a0-4c3e-a26b-be991b26fbf4 · outbound

This paper cites ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.205362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:d20d93f0db68fb2984ebe24f357fa9ec96d1c3548aa1771f2357cfc8ab34ea41

Observation 4cb66a7f-62f9-4bd7-8f8c-178dc392acd6 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:31:11.828360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:290ae7e4cd1e330ed2f56e8fa204180c4f504cb997c8793f983fbc288e5f8430

Observation 8d83bff7-4fb2-40ed-9611-8d9cfe24d9be · outbound

This paper cites LLaV A-Grounding: Grounded Visual Chat with Large Multimodal Models.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition LLaV A-Grounding: Grounded Visual Chat with Large Multimodal Models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.121152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:704bb423fa0fb2214133ad4955504ed60aa43fb469a888447c0ca36c3088e0e0

Observation 19e2f6b4-8eed-4640-bfb3-39f22eacf6fe · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.126405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:f3fddf76552c0860f979cf06e3ed7bee9acee5721caf6d45992fdec0efd74c53

Observation 5536d927-3cea-48a6-84b4-dc3db53c6d52 · outbound

This paper cites Video instruction tuning with synthetic data.TMLR.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Video instruction tuning with synthetic data.TMLR

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.123592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:d87c852273a03adb363b23454beac607802b87a6b93750a115e661afe927fae9

Observation 5b809598-985a-4b88-b931-2654ae26cb36 · outbound

This paper cites Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role Labeling.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role Labeling

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.147600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:922f1f71053ca3dd266dc762d1547bb609c7009a06b881a053af99ff33eece82

Observation 74a752a7-0e9b-4890-8f87-9a198f978264 · outbound

This paper cites Temporal Action Detection with 11 Structured Segment Networks.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Temporal Action Detection with 11 Structured Segment Networks

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.186159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:2e06a2a35a6bf8309f7041f3f30c6256502346391729ba6f388db124030ef171

Observation 3793d8cc-4406-40a6-adb1-e951abd084ea · outbound

This paper cites Streaming Dense Video Captioning.

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition Streaming Dense Video Captioning

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T16:38:03.140877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:50:27.871886Z digest=sha256:aec5e0c91821ef15d072492f8c3780ee98665a01a279bce58c6af7da1b66595f

Pith citing papers

No inbound Pith citation observations are available.