Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:13:21.347468Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2608.08315.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:13:21.347468Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f4370455-d3a7-44e6-8ffb-555f24acd836 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Qwen2.5- vl technical report, 2025
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 146a1edd-ebc8-47b1-a4ea-c5942015e36f · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a9cfc4a-82d4-496a-98c0-3b559f30bb01 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c4f7b43-953f-41cb-b3d1-2fd87e90130e · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Internvl: Scal- ing up vision foundation models and aligning for generic visual-linguistic tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8c9ad2b7-a9c9-4fa9-8386-d75a6fd866a8 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Boundary-aware temporal dy- namic pseudo-supervision pairs generation for zero- shot natural language video localization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 25ff8519-019a-4240-b7d4-79f9b0359976 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No LLM4VG: Large Language Models Evaluation for Video Grounding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3d808a60-3431-49c2-96cd-e5e1dc81cd1d · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Tall: Temporal activity localization via lan- guage query
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0f191f38-52b1-4434-b46c-e62bf6de952e · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No TRACE: Temporal Grounding Video LLM via Causal Event Modeling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5422887-5c32-4b3b-8138-2da02352894b · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video tempo- ral grounding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b8136e6d-ccd2-421c-9180-22b306431425 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Vtimellm: Empower llm to grasp video moments
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e05c93c2-24b4-493e-a89a-b9ec01d698c6 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Lita: Language instructed temporal- localization assistant
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b9058860-27a5-40de-84f5-1693d13c0303 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Granalign: Granularity-aware alignment framework for zero-shot video moment retrieval
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c1368278-c94c-4acb-91a6-878ca53e7f30 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Dense-captioning events in videos
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 174a74fd-7cd0-4e71-b284-ec27f283b44b · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c462f85-4c67-4865-8d8a-f8b8a7f655ab · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Detecting moments and highlights in videos via natural language queries.Advances in Neural Information Processing Systems, 34:11846–11858, 2021
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b8e0bcb6-93c9-40c2-8b3f-4bfb5f079b1c · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec53043a-a89c-4557-88d9-a23810dec103 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Evaluating object hallucina- tion in large vision-language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c13c99db-45be-439b-9235-2d83cba45546 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Univer- sal video temporal grounding with generative multi- modal large language models.Advances in Neu- ral Information Processing Systems, 38:64426–64455,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cceff855-71e5-4e6f-ba9a-1e0a21bb342e · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Univtg: Towards unified video-language temporal grounding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c5d0fda9-42bf-4b91-8434-d9b5c817af67 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Enrich and detect: Video temporal ground- ing with multimodal llms
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1b6e6aad-a810-4233-808d-653eef6b5a3e · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54dbf097-00f9-4d99-931e-a7969e1afc02 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Chatvtg: Video temporal grounding via chat with video dialogue large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8887fdb5-cfa2-4fda-a68d-ecd742e93200 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Grounding action descriptions in videos
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1f43cfb5-22ef-4a9d-b78e-a86fdb398536 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2f283da1-a739-4f44-bc42-ef5ab8393b3d · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No HawkEye: Training Video-Text LLMs for Grounding Text in Videos
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4db2ec45-81ae-4ab3-8a3e-da7b3f6d7a18 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Time-r1: Post- training large vision language model for temporal video grounding.Advances in Neural Information Processing Systems, 38:83330–83364, 2026
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 424deaef-f576-41d1-a14b-3f8be6894ee3 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Negative sample matters: A renais- sance of metric learning for temporal grounding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 459ba1b0-6648-4086-b995-81b0cb4d3965 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Zero-shot video moment retrieval via off-the-shelf multimodal large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d00a842c-ee51-438b-bfdf-71fdd08c529f · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No mplug-owl3: Towards long image-sequence under- standing in multi-modal large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f34c7a9e-e391-49f7-8284-507e939154c6 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Time- suite: Improving mllms for long video understanding via grounded tuning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1b962534-b04a-4556-9197-c589b18447ef · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Learning 2d temporal adjacent networks for moment localization with natural language
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2a6ce923-819b-48ab-8fa0-aec4c9694cec · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Llava-next: A strong zero-shot video under- standing model, 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e3804e84-edea-41ef-8d6e-478db517e334 · outbound
Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No Omnivtg: A large-scale dataset and training paradigm for open-world video tempo- ral grounding
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
No inbound Pith citation observations are available.