Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:50:12.188418Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2504.14553.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:50:12.188418Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ff6acc08-898c-4ae0-9029-d57bb21daa57 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Localizing mo- ments in video with natural language
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0e13720-1928-45a9-86dd-75ba2afe9f44 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Boundary content graph neural network for temporal action proposal generation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation df77c5e4-7a07-454c-b8a8-f91ac01b52b7 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08cea097-311e-41dc-a647-abdd6cc434c3 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Activitynet: A large-scale video benchmark for human activity understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 658e8f35-3558-4537-9c3a-8055ef4a7af5 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3c77918f-81cc-4e1d-8347-09e1f7b829a8 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Tallformer: Temporal ac- tion localization with a long-memory transformer
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e11e8444-c6d8-4973-ae27-f74025a37496 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Vindlu: A recipe for ef- fective video-and-language pretraining
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cf8286c5-9cd7-4431-b86b-0ea317814ccb · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Bert: Pre-training of deep bidirectional trans- formers for language understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d0402a72-4fe6-4d3e-b4ee-84a64b963430 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection End-to-end learning of motion representation for video understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b6cf7a-9704-4c7d-8a74-8db79551633c · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Tall: Temporal activity localization via language query
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 660646cb-6379-487f-be8e-4bbb50558477 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection in the wild
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9609bb16-1aa7-416f-820d-7f0bd7386e1b · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection T-rex2: Towards generic object detec- tion via text-visual prompt synergy
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 084166b9-93db-4129-8882-125be8f4183a · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Prompting visual-language models for efficient video understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 178afd0a-52a8-4b11-b021-c28885dd72f2 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Dense-captioning events in videos
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4ff0dbfe-99f5-4758-a122-9aeb10c15eb2 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection VideoChat: Chat-Centric Video Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 858e4a86-03dd-4367-9de3-edb0737d1eba · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Unmasked teacher: Towards training-efficient video foundation models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2a6bc45b-887f-407f-b02c-6b591c815850 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dddf181c-cfef-475b-a87c-eae5063a370a · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Grounded language-image pre-training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e2326f9c-a21e-4f9c-9d64-796a67edefe8 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Detal: open-vocabulary temporal action lo- calization with decoupled networks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6a64592b-ffe5-4217-89ae-bc456465e19a · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Learning salient boundary feature for anchor- free temporal action localization
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 357ce09c-b620-473f-bb54-4355dd7f0a32 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Tsm: Temporal shift module for efficient video understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e3ed86d-c28e-45fb-9936-2273393539ee · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Univtg: Towards unified video- language temporal grounding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2347ab74-c07c-4f6c-84f6-0ae9820ffe81 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Single shot tempo- ral action detection
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8656353c-c7b3-4b6d-a36d-4f8a7e93d3c0 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Bmn: Boundary-matching network for temporal action pro- posal generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47621ad1-fd99-4c7e-b9f6-bb4c2a3d2bfc · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Focal loss for dense object detection
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf5fdd36-f2d4-4aff-8841-0fe11f45a718 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3b4eb629-b65c-4e19-ae26-6ee52a125700 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection End-to-end temporal action detection with 1b parameters across 1000 frames
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cd97fa05-a8d4-49d2-9f87-c17f7b56e8a6 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection An empirical study of end-to-end temporal action detection
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 18234106-3f50-4670-81aa-8f358e43d1a9 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 38ffb9e2-8104-4f0b-9a22-8f88d600f0ff · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32f8dd08-93bd-472f-b755-fbb32083ec25 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Fineaction: A fine-grained video dataset for temporal action localization
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 39fbbbe0-5869-48c1-96d7-3d761f13c087 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Decoupled Weight Decay Regularization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c560b90d-9e59-4f4c-804b-1234c681e8f8 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2e37c068-59d9-4c0f-9b1d-cfd7baf896a2 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Valley: Video Assistant with Large Language model Enhanced abilitY
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93bcd896-c203-4d59-ab08-f32a11c4296b · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Local- global video-text interactions for temporal grounding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccc8222a-f8fc-4507-aea1-9c0035738b10 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Zero-shot temporal action detection via vision-language prompting
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 53bcc1ca-7d58-498b-8e3e-6180ca02c5f5 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 704614e2-b4dd-45cc-ab92-4d7fb5de9683 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Learning transferable visual models from natural language supervi- sion
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee6e1fd8-a83c-4c4d-aea8-76e3142a7f0a · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Action sensitivity learning for temporal action localization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aca31db7-1f00-475c-9d08-e9018784dd56 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation db9fe65e-e596-4c62-bebd-c5361b85c16d · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Tridet: Temporal action detection with relative boundary modeling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8772812d-fe1e-4bc0-a66a-2b73c5e46dd3 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 941ccdad-9b86-4eaf-b617-de3919e7f0cc · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Attention is all you need
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 537cd032-6373-483a-825c-aab3c70f68b2 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection ActionCLIP: A New Paradigm for Video Action Recognition
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d8d9532-581a-4436-92de-8396b86cd79a · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 739c4d95-ae11-437e-ab3d-0315a537f523 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9679374f-8457-46e5-a818-bda04611e9a7 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Learning to refactor action and co-occurrence fea- tures for temporal action localization
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ae7d4550-78fd-452f-b34e-9ac9f0ff334f · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Unloc: A unified framework for video localiza- tion tasks
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 900d7373-058c-4c5a-84e0-1a4d3b8d4034 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Basictad: an astounding rgb-only baseline for tem- poral action detection
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dbaeb6b5-e803-49a0-bd41-e7ae532faca2 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Detclipv3: To- wards versatile generative open-vocabulary object detection
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fe123d0a-b629-4d54-9574-a435d9734cd6 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Semantic conditioned dynamic modulation for tempo- ral sentence grounding in videos
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bfbcb11-769e-4293-8ec2-9b0570a481b5 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Graph con- volutional networks for temporal action localization
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e8639064-53de-4d74-991a-83d9d2e4c7cb · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Dense regression network for video grounding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df438b76-7cb8-466f-9a6c-b5bc4ccfbe76 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Unimd: Towards unifying moment retrieval and temporal ac- tion detection
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 982e96c6-3a4d-47e7-b8f3-56993819cb5b · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Actionformer: Lo- calizing moments of actions with transformers
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5c32f57-678f-4876-9b73-922107083c3f · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03f4f722-528f-4aef-af00-43f5aabe29fa · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Learning 2d temporal adjacent networks for moment local- ization with natural language
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 60f187a1-fe2d-4004-abd5-18bdc29a4533 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Hacs: Human action clips and segments dataset for recognition and temporal localization
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d25f12e9-52ae-4d3b-b006-57dde7b74ee4 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Distance-iou loss: Faster and bet- ter learning for bounding box regression
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a39519a1-9e4f-4c62-8deb-c1b9232acb08 · outbound
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Enriching local and global contexts for temporal action localization
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
No inbound Pith citation observations are available.