Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:19:40.237501Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2507.19599.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:19:40.237501Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
80 of 80 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f4a2b7bb-2e0c-4a06-b9fb-9f8b22e30d21 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a59f5c-3d8f-4f09-ab56-e0507cb122af · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Burst: A benchmark for unifying ob- ject recognition, segmentation and tracking in video
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 72cd616b-060b-4990-93c0-ea6b3f62f79b · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64102306-6b7f-4f91-80c3-f32aa1652cfa · outbound
Object-centric Video Question Answering with Visual Grounding and Referring One token to seg them all: Lan- guage instructed reasoning segmentation in videos
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation de2d2fdc-fc69-4e84-a048-286ac09946bf · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Xmem++: Production-level video segmentation from few annotated frames
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3897ae52-368e-4cf8-a23e-9833ac40268c · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Coco-stuff: Thing and stuff classes in context
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6fca08f5-55d2-42a4-8ad2-c38d9365bb64 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Vip-llava: Making large multi- modal models understand arbitrary visual prompts
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5d06ad8b-31d9-4502-86fc-bb077eb6f79d · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ee91e18-9a87-407e-9fd5-9213426db69f · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8d0e2f8-5dd9-4fac-8c69-b65e8e9c55ea · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Detect what you can: Detecting and representing objects using holistic models and body parts
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87cfd5da-20c0-4dd2-be86-c82accb6bea4 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08867776-fa48-4a4a-afc0-7032580eddf9 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 467cceab-2790-4e27-8aab-b4c40e40fda3 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e4f8a82-6789-47d6-ab6f-b8fe37747739 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Grounded question- answering in long egocentric videos
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b9e99d39-f6dc-4276-9c58-823895421b22 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Mevis: A large-scale bench- mark for video segmentation with motion expressions
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d3ab6386-277c-40ad-a059-8b69694a63cc · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Mose: A new dataset for video object segmentation in complex scenes
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a69ad915-4348-4f40-b026-8ab69739055a · outbound
Object-centric Video Question Answering with Visual Grounding and Referring LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 615ca1a8-b839-481e-86be-2a949d326ed0 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring LoRA: Low-Rank Adaptation of Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a6d0135-259a-4498-af7e-5c2be488fbd3 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring GPT-4o System Card
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4d82382-ce65-4f2f-90f6-875055bd318e · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Cotracker3: Simpler and better point tracking by pseudo-labelling real videos
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8e024cfc-dbb8-4abf-99fb-9249ed304db2 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Referitgame: Referring to objects in photographs of natural scenes
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 29c9ab82-2837-4706-add4-46a147faccd3 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Video object segmentation with language referring ex- pressions
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a306dbb0-bdd4-4aa2-8e53-6b96dab7dc20 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Segment anything
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 79e37907-115d-4a89-8f14-eeb865c5259e · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Grounding language models to images for multi- modal inputs and outputs
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 58447ff6-88b7-4e0d-ab6b-135f1d9fc163 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Generating images with multimodal language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d3f1ad95-2d6d-41cb-b41c-26b35e07aed9 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Lisa: Reasoning segmentation via large language model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f50cf220-2210-43f9-852c-2d9ad08ed121 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring MIMIC-IT: Multi-Modal In-Context Instruction Tuning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95ea52ed-4eec-4bce-990b-893c9990c20e · outbound
Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-OneVision: Easy Visual Task Transfer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7cad20f-21af-4df1-92e8-4210147ffecf · outbound
Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 172277c2-04cd-4578-9645-535d88e70850 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3a772343-879c-4ae3-8d14-0f5d729134dc · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Towards Robust Referring Video Object Segmentation with Cyclic Relational Consensus
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b453265-6047-4a43-867c-f4127e8ee3d2 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Describe anything: Detailed localized image and video captioning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b81e087f-251f-40a7-ae09-ceb99b331e7a · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1237dd00-ff9e-45c2-adab-a977ca078d9e · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Rouge: A package for automatic eval- uation of summaries
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a155d143-04a0-4678-ac25-5bd7cbdcefed · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Glus: Global-local reasoning unified into a single large language model for video segmentation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a14696b6-30a5-4cce-8ba6-08e350b0cc02 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Visual instruction tuning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 47c75236-174c-43bb-82be-0cc419d7d744 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Lamra: Large multimodal model as your ad- vanced retrieval assistant
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2dc44f75-013a-403d-be4b-9560b923ed50 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Decoupled Weight Decay Regularization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa765c0b-ad33-427e-93f2-d3503289bb97 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cbe70e5-59c8-4e31-98c6-eee78a3e78ec · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Generation and comprehension of unambiguous ob- ject descriptions
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1ea753bb-c75a-4a3c-af41-dcc1cd2756be · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Large-scale video panoptic segmentation in the wild: A benchmark
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 18aacfb3-e0e4-4215-bf55-5ba9bff3d63f · outbound
Object-centric Video Question Answering with Visual Grounding and Referring V-net: Fully convolutional neural networks for volumetric medical image segmentation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0ed6dda5-284f-482d-af20-571d707c01f7 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Bleu: A method for automatic eval- uation of machine translation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4fc8125a-1305-44cd-931f-26d4ddd572c4 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Perception test: A diagnostic bench- mark for multimodal video models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 07e3f2f3-ce0a-4fc7-b345-467de8bfbfab · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa54ea9a-a496-4f7e-aace-b97d167440a3 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Occluded video in- stance segmentation: A benchmark
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c5968d3d-6c7a-4dee-8860-bda3c817aa7e · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Artemis: Towards referential un- derstanding in complex videos
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fc7d881a-c0f5-4e22-b453-0570dc9067a9 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Paco: Parts and attributes of common objects
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ca092a3f-da64-4fc2-aee2-c549055601d9 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring SAM 2: Segment Anything in Images and Videos
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c03be5e-3b00-4980-88c4-78066345f1ff · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Hiera: A hierarchical vision transformer without the bells-and-whistles
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 297c4461-aa4a-413f-8cf5-41104126d5c6 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Urvos: Unified referring video object segmentation network with a large-scale benchmark
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 624e5f94-9f0c-470f-acef-f2167f960c44 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Emu: Generative Pretraining in Multimodality
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b464f74-7deb-44c6-8ad2-d543e84718e8 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Cider: Consensus-based image descrip- tion evaluation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc9be667-cf9b-423a-94e6-5f3465b07219 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Ov-vis: Open-vocabulary video instance segmentation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a4c34e4a-92bc-43f1-8a2c-09516a9bef1c · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff284c6b-ee5d-44f0-b956-a21b58c65482 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Unidentified video objects: A benchmark for dense, open-world segmentation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6a9aaa68-a589-4f06-9676-baa83d4ef10e · outbound
Object-centric Video Question Answering with Visual Grounding and Referring InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 177d8f52-5c82-4d46-86ca-5d8315692384 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Internvideo2: Scaling foun- dation models for multimodal video understanding
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ff67cfcd-c3e3-4eb0-bb9c-114846a0b438 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Next-qa: Next phase of question-answering to explaining temporal actions
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 872524ac-7cfc-420a-a0ea-c0d3541c37e5 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Visa: Reasoning video object segmen- tation via large language models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 455e99f2-4c90-42d3-be2f-6e87c41904f0 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring VideoGPT: Video Generation using VQ-VAE and Transformers
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a2257f-8dea-4cd1-8131-6d0e80f5ee68 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 827c5a9c-da2c-4dd7-80b6-f910464b5d2e · outbound
Object-centric Video Question Answering with Visual Grounding and Referring LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb2f858e-5f53-4f4b-b6dc-fac94a69787c · outbound
Object-centric Video Question Answering with Visual Grounding and Referring mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd125f56-671f-4185-abee-9e1900f90ef0 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 450421f4-78e9-495d-b65b-092441bda764 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 972d9ea1-6637-4f61-8713-1686d8528fd1 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Sa2va: Marrying sam2 with llava for dense grounded understanding of im- ages and videos
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a545401d-1949-49a9-8bc6-1d7869e20b88 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Os- prey: Pixel understanding with visual instruction tun- ing
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ccf18085-1bfe-4023-b3b1-542b44531262 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Videorefer suite: Advancing spatial-temporal object understanding with video llm
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 025328a4-de54-48dd-bb8d-b7f8679de765 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7f6f91c-70a2-46e0-adda-b72c587b9531 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99f4e4f6-d6ea-447e-b0c2-7f198124978f · outbound
Object-centric Video Question Answering with Visual Grounding and Referring LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5ee245e-d564-4974-a412-7de86e962e88 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67cb2719-8b81-48f7-adeb-ca73373abd9c · outbound
Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d770155-febd-4a35-a8c6-8d9ae90f717e · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Scene parsing through ade20k dataset
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f836cefc-0b8a-427a-9a92-f9483ff2bffc · outbound
Object-centric Video Question Answering with Visual Grounding and Referring MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b64863f9-c6e2-441b-8d25-9f17e8265ac1 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring The complete list of used datasets in training is presented in Tab
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 129b29ba-459e-434b-8336-2757de218c6a · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b5b5aa3e-f768-44d4-b4c0-9b6077d8db28 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c21f2be1-ac58-4b45-9031-bf31ee57b6f1 · outbound
Object-centric Video Question Answering with Visual Grounding and Referring Unresolved cited work
Reference 223
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.