Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T00:54:53.789523Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 0 inbound Pith citation observations for arXiv:2512.06673.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T00:54:53.789523Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
86 of 86 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d39bed59-f4bd-46df-b1b2-9bcfe84c731f · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Localizing mo- ments in video with natural language
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 14512fa4-c8ed-4176-b7b8-789f16e95ff5 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Lawrence Zitnick, and Devi Parikh
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ae986ee4-75aa-4e6f-ab89-be88cdb5fd34 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c8621c7-08a1-48c5-8f1a-20f15b7302f8 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding One token to seg them all: Language instructed reasoning seg- mentation in videos.Advances in Neural Information Pro- cessing Systems (NeurIPS), 37:6833–6859
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81260634-4bc2-4b18-819a-4d363323e5ca · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding End-to- end object detection with transformers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 17b636cd-f3b3-4751-870c-f9954cd11183 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Chen and William B
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 91bdf40b-93f9-4bc6-8bdb-ad0bc3fe7364 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 559cd5a6-674d-4b9a-85c5-f993e80f5adc · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Sutrack: Towards simple and unified single object tracking
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1b18494c-ce0e-40c8-a5a4-ed71a1834bc6 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 478d5077-302c-4dfc-bc15-5b34516c64e8 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 05787ab9-47af-4202-a225-2e5012a2b0f3 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ea1b4a1f-8dee-4cb9-af8a-f60352e59c7a · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 605e6d0c-d22d-473b-81fc-5adda7a1f7b7 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04a5bd31-04a1-4392-ad73-e48d61960717 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c70e0ede-58dc-471e-b6dc-af7f568b3217 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Tall: Temporal activity localization via language query
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d92ec788-1196-4346-9dd3-47634f82f031 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dda6359b-8f55-455f-bebd-98c639fb5b2a · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Ego4d: Around the world in 3,000 hours of egocentric video
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b2deb3ce-b3e1-4338-b375-6b3edea312e9 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Agqa: A benchmark for compositional spatio-temporal reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b8ba7102-3c03-4912-bf30-b8486ea8f632 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Context-guided spatio-temporal video grounding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 62678143-30fe-4a46-942c-f04e3065bf57 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f7a8c12-6772-493b-afab-5c4d15ca4c88 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding TRACE: Temporal Grounding Video LLM via Causal Event Modeling
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9cfe936b-f2db-4cb4-9de6-098390cdaf1e · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fa5dab23-0308-4fde-b04b-3d7f4835fb54 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Creating summaries from user videos
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f63be08b-25b4-4032-b6ef-d97e7fbbaa0c · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Lora: Low-rank adaptation of large language models.Inter- national Conference on Learning Representations (ICLR), 1 (2):3
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb000905-2dd9-4bda-ac22-47530f9a5e60 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Vtimellm: Empower llm to grasp video moments
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e31a8702-54d6-4968-be8d-f0a0d2829b17 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Lita: Language instructed temporal-localization assistant
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 289f36d7-8a6d-40d6-9866-7730ef4ab9a7 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding GPT-4o System Card
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9c37f211-c121-497d-be4f-965d2a61ffbb · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Embracing con- sistency: A one-stage approach for spatio-temporal video grounding.Advances in Neural Information Processing Sys- tems (NeurIPS), 35
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ad30844a-01db-414c-9e27-a4ba1cf84b14 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Language repository for long video understanding
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation abdc1bd4-909d-4223-b571-bcd3f492cd85 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Referitgame: Referring to objects in pho- tographs of natural scenes
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c7763273-4139-44aa-95a3-3aef49417473 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Dense-captioning events in videos
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d3550782-3e28-4ce4-9811-8939a0bc7ba8 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Berg, and Mohit Bansal
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f1ed475d-ccd7-4a0d-8357-fa598ae586fe · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Detecting moments and highlights in videos via natural language queries.Advances in Neural Information Processing Systems (NeurIPS), 34:11846–11858
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c75648d-0ef4-49e2-b309-a9a3f18588ce · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 44b6eeb1-b677-4c82-a566-fdfe4a14c356 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Llava-ST: A multimodal large language model for fine-grained spatial- temporal understanding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c26515f-7644-4eaa-8667-6f0379f61cb5 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Llava-st: A multimodal large language model for fine-grained spatial- temporal understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4926f75f-0dfc-422e-8ca1-4d41c72e0e5a · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding VideoChat: Chat-Centric Video Understanding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dab7d110-72cf-45aa-bde7-ede966d4b50a · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 98d8840d-f323-4d65-b85c-e426373fbeed · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Referdino: Referring video object segmentation with visual grounding founda- tions
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 055628cf-ae76-45e4-b59b-f1189c9bdaf5 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Video-llava: Learning united visual repre- sentation by alignment before projection
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 18a6556e-e704-41d1-8a9f-e41e1cc07152 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Glus: Global-local reasoning unified into a single large lan- guage model for video segmentation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 73071e62-ca17-42ba-9f8e-93748517271f · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Collaborative static and dynamic vision-language streams for spatio-temporal video ground- ing
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e70e5bf2-40dd-446c-b33f-21eb56bef62c · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Grounding DINO: Marrying dino with grounded pre-training for open-set object detection
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5989b7f1-3842-456a-84c9-24deb37c597d · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding TempCompass: Do Video LLMs Really Understand Videos?
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 854c8ae5-24f8-4e62-983c-a1376105faee · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Swin transformer: Hierarchical vision transformer using shifted windows
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 470f744a-a41c-44e2-9102-190dfa7b92dd · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3780dad6-ab2a-414b-8f9d-9e81fe75ccdc · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ad8f6c2f-e4f5-4540-b780-e9d6848fd243 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Video-chatgpt: Towards detailed video un- derstanding via large vision and language models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 91e44e5a-74f7-4fa7-8e38-f613afa1264f · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Genera- tion and comprehension of unambiguous object descriptions
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc73d32e-0f37-46d9-acae-47524f284c9b · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6a904aa3-5c33-48a0-a3d9-cf921f28d043 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f98fbac6-d048-4c40-aa82-6c942f6f8cea · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Streaming long video understanding with large language models.Advances in Neu- ral Information Processing Systems (NeurIPS), 37:119336– 119360
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 24f6e6ea-3ca1-41f3-ac95-1482a9a73e57 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding SAM 2: Segment Anything in Images and Videos
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e07ad01-53fd-4060-9769-0c15e0ec069f · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Ground- ing action descriptions in videos.Transactions of the Associ- ation for Computational Linguistics (TACL), 1:25–36
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0bbdfc4e-8885-4c6e-bfd0-55933f105876 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e5eac96-d3a7-4977-b03f-6de7d9092c6a · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7ec0f3a3-2e73-4831-8b4c-eec4f2f10c1e · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Moviechat: From dense token to sparse memory for long video understanding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e85449fd-1b50-4556-9d36-f763e20685f9 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Tvsum: Summarizing web videos using titles
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 60eb5f8a-2ab5-4393-8d2f-8521301de291 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding STVGBERT: A visual- linguistic transformer based framework for spatio-temporal video grounding
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dc50b21c-4e74-489d-9e1e-b41505fdee92 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Human-centric spatio-temporal video grounding with visual transformers
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 19a5d37e-6486-4a8a-b153-7e6e0d905da5 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 994323ee-fe0b-49b7-8ad5-1f43387f8019 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 040a74d2-2dee-4802-b188-0c456b63ffee · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 39b8fe18-aaaf-4c37-9a40-2fc6a16bc31c · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Can i trust your answer? video question answering
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 26411682-4ebb-4120-b129-3efad12db353 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Msr-vtt: A large video description dataset for bridging video and language
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5ce41bf4-68fd-4d56-847f-814f054843f0 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Visa: Reasoning video object segmentation via large lan- guage models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2cd3a306-6a8e-474e-96b0-bb55670525d1 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Task preference optimization: Improving mul- timodal large language models with vision task alignment
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93b6b99a-001c-48a6-b010-fd366be95e5f · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Tubedetr: Spatio-temporal video ground- ing with transformers
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2192f487-8071-427c-b643-8aa91420d2df · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Zero-shot video question answering via frozen bidirectional language models.Advances in Neural Information Processing Systems (NeurIPS)
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 79f23c7e-a294-4381-830c-ae8e0993b51a · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Qwen3 Technical Report
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3487de46-2ea9-4107-9fa9-66e23515b09e · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Tenenbaum
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a8fb2e04-5606-460b-a560-f9ee8b8dc6b6 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Self-chained image-language model for video localization and question answering
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0bf8e50c-39f6-4e84-860c-d0531d8dfc4e · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d3f6849d-288b-4d9f-86fd-6668e103d25e · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Sigmoid loss for language image pre-training
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 645d8179-2cb5-4179-9228-8250a797ca45 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b9a1aa6-493e-42f0-baad-440c81e1a141 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding A simple llm framework for long-range video question-answering
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f4f65383-30d6-4e85-a504-81211bb2062d · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 129f71a8-9fbb-4def-97d4-53c46bd94599 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 890a4901-07c9-41c8-bfd3-773d0ae1204e · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Where does it exist: Spatio-temporal video grounding for multi-form sentences
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5c190997-ec5f-41ba-b72f-d5d5304883e6 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding An Open and Comprehensive Pipeline for Unified Object Grounding and Detection
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 733014c9-ed2f-488a-b7dd-faff8d6fd368 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Given textual featuresQand video featuresV, LLMs employ a textualized box decoder to predict a sequence of tokensy 1:L = (y 1
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b75207c1-56e7-42a1-9fba-a15237e2014c · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding <image> Locate the visual content described by the query <query></query> in the image
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3dffa583-cbad-45d7-b506-2eaafeef2ccf · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding What‘sundertheskier’sfeet?
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 53cd6288-2f35-4907-8d4b-8d88dd258ac2 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Rows marked with*are zero- shot settings
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 791fcf43-268b-41ec-97f0-4e99d341ef3a · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding 6–11 present additional qualitative results across im- ages and videos
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2f563edb-f2f8-474d-a3e0-9305b23cfa48 · outbound
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Consequently, DEViL is trained to emit a single RST and retrieve one tube per query, without explicitly modeling multiple entities or their roles
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.