Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-01T08:41:31.700367Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2604.25886.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-01T08:41:31.700367Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 140600bd-b3cc-4483-88b0-50acf0d6d1a7 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Univtg: Towards unified video-language temporal grounding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 641867e7-b10f-4d3c-85b0-b29f92d5e7ed · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Context-aware biaffine localizing network for temporal sentence grounding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dcf9ca09-3fc6-4e26-8af9-3d6855a4d2f3 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Internvl: Scaling up vision foun- dation models and aligning for generic visual-linguistic tasks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c3f5d603-35ae-48d7-a3ee-45be3ec592d0 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Chatvtg: Video temporal grounding via chat with video dialogue large language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c33fb6a7-93b8-410b-ab79-b58bbd373e39 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f81d4d1d-90c9-4b7f-b36d-96c878701b03 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c36caa39-d471-4c62-b1ff-4fbdbf1e6be9 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding On the consistency of video large language models in temporal comprehension
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 21bb46de-5cb1-4227-bbed-d1a16fe0d7e2 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Bridging the gap: A unified video comprehension framework for moment retrieval and highlight detection
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a90dd20a-34cf-41a0-a0e1-d0642cc72b6e · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c2e25f4e-d662-4fb6-9e07-2e55f82db39b · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding LLaVA-OneVision: Easy Visual Task Transfer
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d61d4b4f-e110-4244-8274-137c203bae06 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Zero-shot video moment retrieval via off-the-shelf multimodal large language models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8032a6fe-45fa-4386-82b3-ab100cc10788 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Background-aware moment detection for video moment retrieval
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 965e3597-bd40-45ac-a1c0-7bcbec4c4cac · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Dense video captioning: A survey of techniques, datasets and evaluation protocols
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f8718ead-6a14-44fa-9856-a073aa767d69 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Dense video captioning using unsupervised semantic information
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a26210ed-34d9-437f-a76d-b29c279b15a3 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Lvd-2m: A long-take video dataset with temporally dense captions
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eba31263-66a4-4c5a-83fa-12ba049f3164 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Unsupervised video highlight detection by learning from audio and visual recurrence
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f1752f48-a684-4897-a918-b42a02f1af8d · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Less is more: Learning highlight detection from video duration
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f69f8d2e-543c-4534-9bb4-8d5d1e5d34fa · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Contrastive learn- ing for unsupervised video highlight detection
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6815b9c0-d40d-49b5-9b76-deabc2207972 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Query-dependent video representation for moment retrieval and highlight detection
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 97f42890-b4db-45fd-936a-b333dc3cdd4d · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Tvqa+: Spatio-temporal grounding for video question answering
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7610431d-2804-432a-ba4d-cfd0c90a941d · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Timecraft: Navigate weakly-supervised temporal grounded video question answering via bi-directional reasoning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 063e4e36-3f3d-4609-889c-187dfc7381af · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Grounded question-answering in long egocentric videos
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7dc83e2b-e036-4bdf-99a3-5ca86ee44eb5 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Can i trust your answer? visually grounded video question answering
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5be6577f-0360-4df6-bcda-85ac9fc5504d · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding A survey on video temporal grounding with multimodal large lan- guage model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d1ffa588-fc59-40b6-9d71-a39139421e77 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d6a29860-7a1a-422e-9de5-4fb838b158e3 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Alvarez, Lei Zhang, and Zhiding Yu
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 90111710-33c5-404b-9afb-80e8eb15db3c · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Towards visual-prompt temporal answer grounding in instructional video
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4ffc359c-2fbf-4611-a7e5-00a4a5286d97 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Number it: Temporal grounding videos like flipping manga
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 03fad860-5072-48ff-bc0f-6a89113d0702 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Vtimellm: Empower llm to grasp video moments
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 54bd5be4-ee64-4b75-afce-bad464169b8b · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 37509041-2377-4856-9312-547eb491ec15 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Training-free video temporal grounding using large-scale pre-trained models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 055d4cd6-6130-414a-b394-1959201a0180 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding TRACE: Temporal Grounding Video LLM via Causal Event Modeling
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 439a5904-bd81-4259-9ab6-974e0356269a · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Omni-rgpt: Unifying image and video region-level understanding via token marks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 67edd1eb-495a-4cd3-8238-ec895844a8ee · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 26f41eed-5359-4dcb-a9f4-e9fd5ac1600b · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Sa2va- i: Improving sa2va results with consistent training and inference
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1e33ad75-ace6-4194-9d3c-7c060947b04c · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Videoglamm: A large multimodal model for pixel-level visual grounding in videos
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7bbc80e7-0eeb-43aa-aea5-a9a56bb47d97 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Videorefer suite: Advancing spatial- temporal object understanding with video llm
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d8e9173-957b-45ba-a782-b4599bbda1dc · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Vip-llava: Making large multimodal models understand arbitrary visual prompts
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c7421cf-7e64-4f08-9a93-2e0f1954ba18 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Generalized decoding for pixel, image, and language
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 53a9a5f1-e08c-4ec3-a4d5-e6e3ce76512e · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Gsva: Generalized segmentation via multimodal large language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 23edf9ce-eec1-4996-b474-7cdcb2dfd904 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Glamm: Pixel grounding large multimodal model
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 63c51d81-258c-414b-812d-20b1ea2b92f4 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Lisa: Rea- soning segmentation via large language model
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 575d8c05-c7f7-42c1-a20a-3df06e7a08c6 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding CoLLaVO: Crayon Large Language and Vision mOdel
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f9e505eb-d405-4ced-8f5e-c157b04b840b · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Segment anything
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1dfbf54c-991e-4217-83b0-d42f4c33c253 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding SAM 2: Segment Anything in Images and Videos
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8ae83314-bd9e-4409-bae3-9e23baddcca8 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding GroundingGPT: Language enhanced multi-modal grounding model
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5b05204d-7ba7-46d1-8390-b6a87e3d9a5e · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding LITA: Language Instructed Temporal-Localization Assistant
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4ce1d25a-5e4d-49f4-b27a-7c775f745683 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 53319488-39a1-4251-bfa8-9a8cf9e6612a · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7ba36f90-8d5b-4224-860b-09a8bac66510 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b8b842b8-d96b-4015-8c38-0b713131a6bb · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding VTimeLLM: Empower LLM to Grasp Video Moments
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 39369fc2-1877-4b2d-8361-27fbb69b7a60 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 34063787-d5e0-4c4b-a715-e96d5de010e5 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Hawkeye: Training video-text llms for grounding text in videos
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d46e3d6c-f44f-496b-9a3d-863043f454d6 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding HawkEye: Training Video-Text LLMs for Grounding Text in Videos
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b7feabb6-0a7b-4090-9f69-22bb55cfd53c · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 94957e16-60aa-46c2-8b6f-105f9b7a6672 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Tall: Temporal activity local- ization via language query
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2453830b-a0cd-45fa-af24-12106c294f30 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Activitynet: A large-scale video benchmark for human activity under- standing
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 07afdd43-58a4-41f8-bc00-833bafa767f4 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Detecting moments and highlights in videos via natural language queries
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 441a83b3-bcfa-43fe-9b7c-3c71bebbe125 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Long Context Transfer from Language to Vision
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a101b8b2-1b9b-4349-a689-88ed2a3ae3b1 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Yoloe: Real-time seeing anything
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 256f25d8-18cb-442b-8ccc-d94f234e76c4 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Two men both dressed in athletic gear are standing and talking in an indoor weight lifting gym filled with other equipment
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 169c417a-584e-4927-91a6-1a0b8ec978e7 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Overall trends are consistent with those observed on ActivityNet
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 226ef46f-8d68-447e-b4b3-2ec24e30e022 · outbound
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding Table 6 reports the effect of different color parameteriza- tions on ActivityNet
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.