Pith. sign in

Paper Citation Record · LEDGER

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

As of 11 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2607.17423.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17423 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:04:18.778402Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved58
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4fbf2bb5-3bfd-4746-9479-38aa8d899787 · outbound

This paper cites LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:11.814953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:11.814953Z digest=sha256:2bdab26b8d906dfb3fa30a2592ab0ac5f7b44add9178b519ef67642033de1a8d

Observation e47f7bf5-2291-4a4f-ad49-f82cb3617343 · outbound

This paper cites Qwen3-VL Technical Report.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:11.894973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:11.894973Z digest=sha256:0845dc5a5bf05d27f63c5d1160043ee4aa4254272f6f1eee7a1f9b7372a4a5bf

Observation a43b0ddf-e971-4caa-b524-290c6692693c · outbound

This paper cites Datasets and recipes for video temporal grounding via reinforcement learning, 2025.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Datasets and recipes for video temporal grounding via reinforcement learning, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:11.957891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:11.957891Z digest=sha256:278a071de5bca43c37645cfb67b59781f187e877d0cd9b19e92349466fcdad64

Observation 9d1a2cd6-e9c6-4f93-a0d8-a88ce0caa454 · outbound

This paper cites Molmo2: Open weights and data for vision-language models with video understanding and grounding.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Molmo2: Open weights and data for vision-language models with video understanding and grounding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:12.026393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:12.026393Z digest=sha256:88a7b3b862eff40c137548884dd88e058074d669249128b023c792aeb1b1b0df

Observation 059a5502-af0b-4fde-9539-e30507a0718a · outbound

This paper cites Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507 .06261, 2025.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507 .06261, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:12.116730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:12.116730Z digest=sha256:ccabdf4b84f4f81d33c67d6f0d3b93a9c4115d5121d6c2b196560109b674dfa2

Observation 73e57905-d4a7-41ea-80be-b4f245be7ed3 · outbound

This paper cites Videotg-r1: Boosting video temporal grounding via curriculum reinforcement learning on reflected boundary annotations, 2025.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Videotg-r1: Boosting video temporal grounding via curriculum reinforcement learning on reflected boundary annotations, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:12.181473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:12.181473Z digest=sha256:fbfd329354266608b6fce03b1d89ae72540871c0e10ad26c016ac4a8015dd4f0

Observation 0ec66415-43cf-4903-bb10-85bd008682e5 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models, 2026.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Vlmevalkit: An open-source toolkit for evaluating large multi-modality models, 2026

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:12.271974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:12.271974Z digest=sha256:b253593741c1b3192a3c95b5347814bca9df4e3fe6c65146aac6fd6a340b47f9

Observation aca75ebe-4d89-4aec-a7a2-c0deb2987cfc · outbound

This paper cites Tall: Temporal activity localization via language query.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Tall: Temporal activity localization via language query

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:12.314399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:12.314399Z digest=sha256:5bad394df6c25a1dc9a7db5168b897efc2783247f3714100b578aaa7881da3dc

Observation 00a3de38-b3b8-44c6-9dd0-97df952286d8 · outbound

This paper cites Gemini 3: News and announcements.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Gemini 3: News and announcements

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:12.417300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:12.417300Z digest=sha256:242bc56d5c2cec5834a88034e838ca5c1f49134be0cc1a21785c1119e10e2966

Observation 92430fee-c7b6-472d-b36b-30f9699cc9d4 · outbound

This paper cites Gemini 3.1 Pro model card.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Gemini 3.1 Pro model card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:12.520446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:12.520446Z digest=sha256:b81cb17d8f9adc8f8f526a2bc5943151cb456fe9cc030bbc07ac97ccf03ab92f

Observation 9eab2dcb-7639-4823-ab5c-4f1d9fee1ddc · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Ego4d: Around the world in 3,000 hours of egocentric video

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:12.622844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:12.622844Z digest=sha256:fa42b756f044466215097b1b8183f3123fa4df333df4bdbfe0833b932bc7a113

Observation 139f3aab-6821-4eeb-b534-d7e821f6df95 · outbound

This paper cites Localizing Moments in Video with Natural Language.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Localizing Moments in Video with Natural Language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:12.732459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:12.732459Z digest=sha256:a784c068fdaeba295c713c0e9b04c0c8a9e1b0eabcfc7c5befa18c796a9fec2f

Observation 099d850f-03b4-4c49-a708-09c642f738b6 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Vtimellm: Empower llm to grasp video moments

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:12.882777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:12.882777Z digest=sha256:cf2e6fa9947b85edc2da642b949309275e00567b47d19283f809551f24877120

Observation 8c5403d0-92cf-47d4-993f-5a92696f85b8 · outbound

This paper cites LITA: Language Instructed Temporal-Localization Assistant.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs LITA: Language Instructed Temporal-Localization Assistant

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:13.016062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:13.016062Z digest=sha256:a9dff9599f06c71818c584fde9cc0791345c46724fdb0e61f7007a51f30053fc

Observation 179ba791-00e1-48fa-8e0e-2859a7c0df08 · outbound

This paper cites Dense-captioning events in videos.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Dense-captioning events in videos

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:13.152921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:13.152921Z digest=sha256:c15f7a7c85ac8c7640399f9c1f67e64f1a0147ca1a41f2c273e57b77c2efef86

Observation 223b8135-a465-4576-a9fa-84c7378c026d · outbound

This paper cites TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:13.257343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:13.257343Z digest=sha256:b221ea86e56cfef08f26c37f065442c794b1fc6821e517ec23097b354db7a14e

Observation d2ed9961-3dde-4e99-8705-7f2a2c9d4cfd · outbound

This paper cites Detecting moments and highlights in videos via natural language queries.Advances in Neural Information Processing Systems, 34:11846–11858, 2021.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Detecting moments and highlights in videos via natural language queries.Advances in Neural Information Processing Systems, 34:11846–11858, 2021

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:13.355952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:13.355952Z digest=sha256:c30418b80ec9f9d7b776bdb52d88a4c2ecb17a0dbb8afd3964854b74e9bc3bc2

Observation 2c46cb5e-b4d1-46c7-a68f-ac8bc4892e0c · outbound

This paper cites Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:13.530716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:13.530716Z digest=sha256:15f09b81b04839ce4ae2bb76823195cb80a77fdb6d81cf617265dee9737c2fc8

Observation 10e5ab85-771f-4d47-ae18-fd8c099da7a4 · outbound

This paper cites Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:13.665607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:13.665607Z digest=sha256:6e2f22bd12ae02ca8e287cbee2b725860ba3e2b26f31e3fd6a8d1bf228e93513

Observation d5e42eb8-c67e-4065-8607-1efa4d8c3892 · outbound

This paper cites Qwen3-vl-embedding and qwen3-vl-reranker: A unified framework for state-of-the-art multimodal retrieval and ranking,.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Qwen3-vl-embedding and qwen3-vl-reranker: A unified framework for state-of-the-art multimodal retrieval and ranking,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:13.757299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:13.757299Z digest=sha256:af830aaeea4daa3ded7b5ff0a598f4224a028f0ac58c9413815c35c04a9fe24d

Observation b7669212-e82e-4bbf-b512-faae1e162957 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:14.015999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:14.015999Z digest=sha256:1d90ee4d6a1e8a1a962f1f29a58752e73e9a5f4835cf369039ec8ebfb6f4f52d

Observation bf5e9282-a80f-4f2f-a576-9c0d5aed7c4b · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:14.127600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:14.127600Z digest=sha256:263954dd0a99687ea197b2721b5a02cc8547464f107a0642fc56d6bd6ca3d035

Observation 78f1ef15-8a81-402f-8f95-67f877d1628e · outbound

This paper cites Videochat3: Fully open video mllm for efficient and generalist video understanding, 2026.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Videochat3: Fully open video mllm for efficient and generalist video understanding, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:14.235040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:14.235040Z digest=sha256:1bcb2c60a46c1fc2a105055c8eac692faf362fa191b741b961dceeb6fddf5f88

Observation 8b06fb38-f3d3-4955-92bd-09a38536e1d1 · outbound

This paper cites Reasoning guided embeddings: Leveraging mllm reasoning for improved multimodal retrieval,.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Reasoning guided embeddings: Leveraging mllm reasoning for improved multimodal retrieval,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:14.328488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:14.328488Z digest=sha256:c00d74f7e8ee1c3e8dde0d4712e5251772d6e94a2b548345ad4219bbfd092034

Observation 14c19113-53b7-4a41-8a1d-c6fd6fa5dabf · outbound

This paper cites Museg: Reinforcing video temporal understanding via timestamp-aware multi-segment grounding.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Museg: Reinforcing video temporal understanding via timestamp-aware multi-segment grounding

Reference 25

Resolution
verified exact
doi, observed 2026-08-01T18:08:30.024776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-01T18:04:14.546806Z digest=sha256:89c67beeff0b08a2f0f08c41849291c864b463ebbf3fb95e4055f1667d2d2c69

Observation 7f387542-4d9a-48f4-9f19-bcb55b6705d2 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:14.669524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:14.669524Z digest=sha256:d7e943d1664394e06416014eeea9957fb123eac09eca637e64b625cf52d55f99

Observation 1e0810f7-9469-4c59-894d-39418a7e82a0 · outbound

This paper cites HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:14.772882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:14.772882Z digest=sha256:899ad85c4c234daca9e025e1ddd1a9910bd1991f8f94e28d7f5494a0f5b3ae0e

Observation 29babcf6-c2e5-4b27-a890-9923877ce07f · outbound

This paper cites Marlin-2B: A tiny vlm to extract structured information from videos.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Marlin-2B: A tiny vlm to extract structured information from videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:14.881238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:14.881238Z digest=sha256:37bc006c368a9d9b38abb0649546d045473fed502a66772ad4dad8c6b1360fb5

Observation ada93f3e-93ae-4131-abeb-5012eb861180 · outbound

This paper cites Momentor: Advancing video large language model with fine-grained temporal reasoning,.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Momentor: Advancing video large language model with fine-grained temporal reasoning,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.004582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.004582Z digest=sha256:620c39df3fc08229050872af57a9949db8e1a51d67cca52e89cdf11e5207072b

Observation 41a22f8b-3d8d-46a9-8870-cd9ad523c70b · outbound

This paper cites Grounding action descriptions in videos.Transactions of the Association for Computational Linguistics, 1:25–36, 2013.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Grounding action descriptions in videos.Transactions of the Association for Computational Linguistics, 1:25–36, 2013

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.294809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.294809Z digest=sha256:5ce672a1b2248e825c3dd7869b44e8841b64ecb1572c93309801af9a7ed1490e

Observation b0c52bc0-b812-4f45-9d1b-c2b9ab76835b · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.439960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.439960Z digest=sha256:fc2189af142dc4c274dcadc4dc98c2a251a08c44c4d1f6b1f1cd7c6b7ada1175

Observation 2674f1fc-52b0-4580-82df-00b95c7128ba · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.492873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.492873Z digest=sha256:4a3f90d10a454c2008e91ecb836dd1b1cda1b3e2956bf14dd3f34d4fde08ad1e

Observation 7e78a13a-2ef9-42b9-8914-7e31a81b4bb5 · outbound

This paper cites OpenAI GPT-5 System Card.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs OpenAI GPT-5 System Card

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.558089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.558089Z digest=sha256:fd92420e7b27f970231df063584c649333505c41a5e787501528a2c783e9c229

Observation 69bf28ef-0359-471e-8b25-70504e703138 · outbound

This paper cites Mad: A scalable dataset for language grounding in videos from movie audio descriptions.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Mad: A scalable dataset for language grounding in videos from movie audio descriptions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.667362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.667362Z digest=sha256:75dc2ddf83344b423af625915d0683c56f68eed2f167f2e736584ed1c4e58e09

Observation d9222fbf-c7ac-4b7e-a31d-41179b977905 · outbound

This paper cites COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.752349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.752349Z digest=sha256:f727878edead92f938b55b71ee894ec141008de4a2611b50267e477f5248e12c

Observation 26c1c30c-696d-4072-a9f6-de58c86f7386 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Kimi K2.5: Visual Agentic Intelligence

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.818965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.818965Z digest=sha256:2dd7437178fc2e085fb3cd07f366841090126340b12b0b9a35ea6e82ae504555

Observation 67dcbd89-cb3d-4cd5-b5e4-caec2b66390e · outbound

This paper cites Qwen3.5: Accelerating productivity with native multimodal agents, February 2026.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Qwen3.5: Accelerating productivity with native multimodal agents, February 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.902409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.902409Z digest=sha256:4c8bcc2afe2318235a02c27e85e3f09f76482bfbdc0638e24bef911e3a8c1eb5

Observation 48bf1810-df25-4868-a715-14b523634f5d · outbound

This paper cites Vidi: Large Multimodal Models for Video Understanding and Editing.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Vidi: Large Multimodal Models for Video Understanding and Editing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.986807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.986807Z digest=sha256:429286d2b99da8862ee74314e999d04fc33df0f39ff53ce9eca8246b0b415f6e

Observation 360963f2-5000-4297-9c4a-9b390abc450e · outbound

This paper cites Vidi2: Large multimodal models for video understanding and creation.arXiv preprint arXiv:2511.19529, 2025.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Vidi2: Large multimodal models for video understanding and creation.arXiv preprint arXiv:2511.19529, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:16.079361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:16.079361Z digest=sha256:de5d890b49687cb2da81c605eb1b8a25bf61dd3eb99159390503e817ad013ec8

Observation f62f73a4-15c5-4cb1-bffd-ad293bc142cb · outbound

This paper cites Internvideo-next: Towards general video foundation models without video- text supervision, 2026.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Internvideo-next: Towards general video foundation models without video- text supervision, 2026

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:16.161206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:16.161206Z digest=sha256:7e15390c99a4fb035499ecfb0f9836c3f0ffe48898a6e2eced0a535c39ed15c2

Observation 2c1b856a-8498-4fe4-aeb1-ca770c12920d · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:16.246667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:16.246667Z digest=sha256:fce03d01e208ec261a5f96cc00049b89e85801e520dd363bc96d6c26d1f66e74

Observation 22e2741a-b09a-42a9-bf61-b6707c2e31d2 · outbound

This paper cites A Normalized Gaussian Wasserstein Distance for Tiny Object Detection.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs A Normalized Gaussian Wasserstein Distance for Tiny Object Detection

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:16.360822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:16.360822Z digest=sha256:001686ae61db2f58e978a09cc466862d0a6fc4547553ae1b1e7debe8b5bc60b8

Observation 765c908c-d3a4-47e7-956a-7a492c8ae1e8 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:16.535585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:16.535585Z digest=sha256:4b60f510c1d0c9c37538bd3445a0fb796a3d6950e97c593700276abcd7788cea

Observation 0841f998-78fb-4532-98b2-2e9a50486d46 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:16.679249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:16.679249Z digest=sha256:c498020fe4949862d3e16368537d2be686539b89ce4f48d04904a93967df6edc

Observation 31e83b57-04d5-45c2-93b1-b4bffa41c296 · outbound

This paper cites TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:16.823032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:16.823032Z digest=sha256:f64231fc3892ae70c478d9ee5f63210f8de75b74bebc8832e32f94980e321192

Observation 8c9a65e3-14b0-4b3d-bf4e-54edb7612293 · outbound

This paper cites MiMo-VL Technical Report.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs MiMo-VL Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:17.011960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:17.011960Z digest=sha256:0b6fa7c2b4689bf352f2253695ec8e1276ac6ac59c53ff31e6e1df1cd66076dc

Observation 02c8c4c3-0f6e-4e29-b8cc-cc9cf28e9563 · outbound

This paper cites InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:17.133693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:17.133693Z digest=sha256:0ae5094aff77044ae749c11592266b4a1e08393e30f923394e8eab477b60df4b

Observation d42ee9a0-c770-4fb1-9cee-f1355cd38f45 · outbound

This paper cites Momentseeker: A task-oriented benchmark for long-video moment retrieval.Advances in Neural Information Processing Systems, 38, 2026.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Momentseeker: A task-oriented benchmark for long-video moment retrieval.Advances in Neural Information Processing Systems, 38, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:17.291922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:17.291922Z digest=sha256:827e97c95f9de6834d0663409c0af9d72bf961fcca1fb08629ebb07c1e205529

Observation 945ddabc-103f-452f-b73d-73b697b74dc5 · outbound

This paper cites Tempo-r0: A video-mllm for temporal video grounding through efficient temporal sensing reinforcement learning, 2025.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Tempo-r0: A video-mllm for temporal video grounding through efficient temporal sensing reinforcement learning, 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:17.406060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:17.406060Z digest=sha256:0e0e7ea4164d003117ad331c2e8c5f2c5876a1a192c4ba2e509cf1ab04be3f81

Observation 17b4f512-5a9f-4620-9846-3b6e4d70e7fc · outbound

This paper cites Timesuite: Improving mllms for long video understanding via grounded tuning.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Timesuite: Improving mllms for long video understanding via grounded tuning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:17.594384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:17.594384Z digest=sha256:20dfab66f3c65548b7acb83d5afe4f25b4acfe8d12554c767dafde9072000393

Observation f154b320-e218-433b-92d4-f285507ebcde · outbound

This paper cites Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:17.780478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:17.780478Z digest=sha256:7352799974596360a7fc9c8a5aded9fcc1b104bac576e03672f0420b39f1022b

Observation 3b760438-77c5-456c-8925-7ea70c3781bb · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Video-llama: An instruction-tuned audio-visual language model for video understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:17.948936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:17.948936Z digest=sha256:6be0ea6f827fd0bcef5a28cbff8e6396e8700f0baa1d102761369e963bbc7fec

Observation d8aa01b8-baf8-4172-8f68-599747977ac0 · outbound

This paper cites Timelens: Rethinking video temporal grounding with multimodal llms.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Timelens: Rethinking video temporal grounding with multimodal llms

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:18.065417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:18.065417Z digest=sha256:68f4c60e1fea19d332c5f6fadba95fe3e34561b490c2d18086e06810a8a5654e

Observation 52e74d31-b26b-4bc6-839b-6e6b7971cdc6 · outbound

This paper cites Towards Automatic Learning of Procedures from Web Instructional Videos.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Towards Automatic Learning of Procedures from Web Instructional Videos

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:18.270802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:18.270802Z digest=sha256:21c05491e93ff32fd0fe4fc9701db17871081360dce7f4774941c9e11fda7a05

Observation 2bdd718f-3b54-4b0b-a075-dc4bac81a7fb · outbound

This paper cites Dual DETRs for Multi-Label Temporal Action Detection.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Dual DETRs for Multi-Label Temporal Action Detection

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:18.447131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:18.447131Z digest=sha256:9ff4f04f82bda72fb3bbb291b4225f29caab6f135a5b321eb7a319d3b2f9d8bf

Observation b8aa5391-748b-4283-acd0-900dbb0685a3 · outbound

This paper cites Freeret: Mllms as training-free retrievers, 2026.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Freeret: Mllms as training-free retrievers, 2026

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:18.613378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:18.613378Z digest=sha256:30108567545889014e501f83527fda1a25fbe9db58ef0a453ec44785f3a835a6

Observation e920b1e2-5536-473d-96ae-3b20ea849eed · outbound

This paper cites Cross-task weakly supervised learning from instructional videos.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Cross-task weakly supervised learning from instructional videos

Reference 57

Resolution
malformed identifier
no resolver link, observed 2026-08-01T18:04:18.778402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:18.778402Z digest=sha256:39f1f59618420a3ac2033b4e6b172dfec7c28ad1b88ba5485916b5e1cfa29733

Observation 7897806b-4c5e-406a-9871-050898220aa1 · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:15.140313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:15.140313Z digest=sha256:ad7e45c9545ab4ab77f001f7ca243e45f759ad8c3a445d1cd03803bae0b3f563

Observation 134228c4-3783-43b7-a06e-df1c10c2bbd0 · outbound

This paper cites an unresolved cited work.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:14.449844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:14.449844Z digest=sha256:482cd89b6f639ed5683c8ffdf58a386b79f6ecf6438015dad30f0f5e8420c7d7

Observation d278cc39-9585-4a08-a17d-ded003675447 · outbound

This paper cites Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:13.903541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:13.903541Z digest=sha256:cebd7b6e9c65ad281c63cdde77ba10299bbb557ae4eed8c5e571db2a9f7230c6

Pith citing papers

No inbound Pith citation observations are available.