Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T18:04:18.778402Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2607.17423.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T18:04:18.778402Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4fbf2bb5-3bfd-4746-9479-38aa8d899787 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e47f7bf5-2291-4a4f-ad49-f82cb3617343 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a43b0ddf-e971-4caa-b524-290c6692693c · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Datasets and recipes for video temporal grounding via reinforcement learning, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d1a2cd6-e9c6-4f93-a0d8-a88ce0caa454 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Molmo2: Open weights and data for vision-language models with video understanding and grounding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 059a5502-af0b-4fde-9539-e30507a0718a · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507 .06261, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73e57905-d4a7-41ea-80be-b4f245be7ed3 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Videotg-r1: Boosting video temporal grounding via curriculum reinforcement learning on reflected boundary annotations, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ec66415-43cf-4903-bb10-85bd008682e5 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Vlmevalkit: An open-source toolkit for evaluating large multi-modality models, 2026
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aca75ebe-4d89-4aec-a7a2-c0deb2987cfc · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Tall: Temporal activity localization via language query
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00a3de38-b3b8-44c6-9dd0-97df952286d8 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Gemini 3: News and announcements
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92430fee-c7b6-472d-b36b-30f9699cc9d4 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Gemini 3.1 Pro model card
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eab2dcb-7639-4823-ab5c-4f1d9fee1ddc · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Ego4d: Around the world in 3,000 hours of egocentric video
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 139f3aab-6821-4eeb-b534-d7e821f6df95 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Localizing Moments in Video with Natural Language
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 099d850f-03b4-4c49-a708-09c642f738b6 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Vtimellm: Empower llm to grasp video moments
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c5403d0-92cf-47d4-993f-5a92696f85b8 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs LITA: Language Instructed Temporal-Localization Assistant
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 179ba791-00e1-48fa-8e0e-2859a7c0df08 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Dense-captioning events in videos
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 223b8135-a465-4576-a9fa-84c7378c026d · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2ed9961-3dde-4e99-8705-7f2a2c9d4cfd · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Detecting moments and highlights in videos via natural language queries.Advances in Neural Information Processing Systems, 34:11846–11858, 2021
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c46cb5e-b4d1-46c7-a68f-ac8bc4892e0c · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10e5ab85-771f-4d47-ae18-fd8c099da7a4 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5e42eb8-c67e-4065-8607-1efa4d8c3892 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Qwen3-vl-embedding and qwen3-vl-reranker: A unified framework for state-of-the-art multimodal retrieval and ranking,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7669212-e82e-4bbf-b512-faae1e162957 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf5e9282-a80f-4f2f-a576-9c0d5aed7c4b · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78f1ef15-8a81-402f-8f95-67f877d1628e · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Videochat3: Fully open video mllm for efficient and generalist video understanding, 2026
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b06fb38-f3d3-4955-92bd-09a38536e1d1 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Reasoning guided embeddings: Leveraging mllm reasoning for improved multimodal retrieval,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14c19113-53b7-4a41-8a1d-c6fd6fa5dabf · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Museg: Reinforcing video temporal understanding via timestamp-aware multi-segment grounding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7f387542-4d9a-48f4-9f19-bcb55b6705d2 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e0810f7-9469-4c59-894d-39418a7e82a0 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29babcf6-c2e5-4b27-a890-9923877ce07f · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Marlin-2B: A tiny vlm to extract structured information from videos
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ada93f3e-93ae-4131-abeb-5012eb861180 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Momentor: Advancing video large language model with fine-grained temporal reasoning,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a22f8b-3d8d-46a9-8870-cd9ad523c70b · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Grounding action descriptions in videos.Transactions of the Association for Computational Linguistics, 1:25–36, 2013
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0c52bc0-b812-4f45-9d1b-c2b9ab76835b · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2674f1fc-52b0-4580-82df-00b95c7128ba · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e78a13a-2ef9-42b9-8914-7e31a81b4bb5 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs OpenAI GPT-5 System Card
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69bf28ef-0359-471e-8b25-70504e703138 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Mad: A scalable dataset for language grounding in videos from movie audio descriptions
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9222fbf-c7ac-4b7e-a31d-41179b977905 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c1c30c-696d-4072-a9f6-de58c86f7386 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Kimi K2.5: Visual Agentic Intelligence
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67dcbd89-cb3d-4cd5-b5e4-caec2b66390e · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Qwen3.5: Accelerating productivity with native multimodal agents, February 2026
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48bf1810-df25-4868-a715-14b523634f5d · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Vidi: Large Multimodal Models for Video Understanding and Editing
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 360963f2-5000-4297-9c4a-9b390abc450e · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Vidi2: Large multimodal models for video understanding and creation.arXiv preprint arXiv:2511.19529, 2025
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f62f73a4-15c5-4cb1-bffd-ad293bc142cb · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Internvideo-next: Towards general video foundation models without video- text supervision, 2026
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c1b856a-8498-4fe4-aeb1-ca770c12920d · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22e2741a-b09a-42a9-bf61-b6707c2e31d2 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs A Normalized Gaussian Wasserstein Distance for Tiny Object Detection
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 765c908c-d3a4-47e7-956a-7a492c8ae1e8 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0841f998-78fb-4532-98b2-2e9a50486d46 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e83b57-04d5-45c2-93b1-b4bffa41c296 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c9a65e3-14b0-4b3d-bf4e-54edb7612293 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs MiMo-VL Technical Report
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02c8c4c3-0f6e-4e29-b8cc-cc9cf28e9563 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d42ee9a0-c770-4fb1-9cee-f1355cd38f45 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Momentseeker: A task-oriented benchmark for long-video moment retrieval.Advances in Neural Information Processing Systems, 38, 2026
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 945ddabc-103f-452f-b73d-73b697b74dc5 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Tempo-r0: A video-mllm for temporal video grounding through efficient temporal sensing reinforcement learning, 2025
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b4f512-5a9f-4620-9846-3b6e4d70e7fc · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Timesuite: Improving mllms for long video understanding via grounded tuning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f154b320-e218-433b-92d4-f285507ebcde · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b760438-77c5-456c-8925-7ea70c3781bb · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Video-llama: An instruction-tuned audio-visual language model for video understanding
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8aa01b8-baf8-4172-8f68-599747977ac0 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Timelens: Rethinking video temporal grounding with multimodal llms
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e74d31-b26b-4bc6-839b-6e6b7971cdc6 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Towards Automatic Learning of Procedures from Web Instructional Videos
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bdd718f-3b54-4b0b-a075-dc4bac81a7fb · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Dual DETRs for Multi-Label Temporal Action Detection
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8aa5391-748b-4283-acd0-900dbb0685a3 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Freeret: Mllms as training-free retrievers, 2026
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e920b1e2-5536-473d-96ae-3b20ea849eed · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Cross-task weakly supervised learning from instructional videos
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7897806b-4c5e-406a-9871-050898220aa1 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 134228c4-3783-43b7-a06e-df1c10c2bbd0 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d278cc39-9585-4a08-a17d-ded003675447 · outbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.