Pith. sign in

Paper Citation Record · LEDGER

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection

As of 23 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2504.14553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14553 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:50:12.188418Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff6acc08-898c-4ae0-9029-d57bb21daa57 · outbound

This paper cites Localizing mo- ments in video with natural language.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Localizing mo- ments in video with natural language

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:11.935065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:11.935065Z digest=sha256:9482e963ae7f46a790e0622523bfb45a75e7cb6205f349876af40b257b0e3892

Observation f0e13720-1928-45a9-86dd-75ba2afe9f44 · outbound

This paper cites Boundary content graph neural network for temporal action proposal generation.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Boundary content graph neural network for temporal action proposal generation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.940912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:11.940177Z digest=sha256:3d115a570832593619cdb03923c12ed0f98dbd8dadaf7cf67dda104f028352bd

Observation df77c5e4-7a07-454c-b8a8-f91ac01b52b7 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:11.944566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:11.944566Z digest=sha256:66edca63d1454bdbda4e0137f679fbf8018f9716d6f034f58e2c126840611c09

Observation 08cea097-311e-41dc-a647-abdd6cc434c3 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Activitynet: A large-scale video benchmark for human activity understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.917568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:11.948975Z digest=sha256:d3977f6a7a9c836db1e66bb902a9f961da5b7e7e72c3cfeae7fc4aa3c6c6e67a

Observation 658e8f35-3558-4537-9c3a-8055ef4a7af5 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.904327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:11.953584Z digest=sha256:ad7624f00633df3854e50dc72050d4d1ea2c1c438c472a616a347e00114d099a

Observation 3c77918f-81cc-4e1d-8347-09e1f7b829a8 · outbound

This paper cites Tallformer: Temporal ac- tion localization with a long-memory transformer.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Tallformer: Temporal ac- tion localization with a long-memory transformer

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.890742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:11.958554Z digest=sha256:59102d66e6a1a8882c025fdef0e30cfaee72971785775e453f0cda4f803ccdb0

Observation e11e8444-c6d8-4973-ae27-f74025a37496 · outbound

This paper cites Vindlu: A recipe for ef- fective video-and-language pretraining.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Vindlu: A recipe for ef- fective video-and-language pretraining

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.877658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:11.963159Z digest=sha256:e523cc326c9b3c51ac0c52a31df15637e87c4fd2a3255183366cf0b0dc11be7f

Observation cf8286c5-9cd7-4431-b86b-0ea317814ccb · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.864111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:11.967899Z digest=sha256:a08a1328fbfcebe7258363eccca04f24d388980c80e4f2a50abdc2f57f766a72

Observation d0402a72-4fe6-4d3e-b4ee-84a64b963430 · outbound

This paper cites End-to-end learning of motion representation for video understanding.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection End-to-end learning of motion representation for video understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:11.971950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:11.971950Z digest=sha256:98f95593e395da5b87924c7356eb893393ab7ac39b83ce32c7d7564430e8c161

Observation 07b6cf7a-9704-4c7d-8a74-8db79551633c · outbound

This paper cites Tall: Temporal activity localization via language query.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Tall: Temporal activity localization via language query

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.840209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:11.976006Z digest=sha256:21c67a752323fdc45252f88ffb5e1bf59b4f1fb8e6d2f4795ab15845a5f83598

Observation 660646cb-6379-487f-be8e-4bbb50558477 · outbound

This paper cites in the wild.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection in the wild

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.825150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:11.980125Z digest=sha256:4fa4d7582088442fd04fb7455db4fd8132b00412b6e2c93413b07a2634b5218d

Observation 9609bb16-1aa7-416f-820d-7f0bd7386e1b · outbound

This paper cites T-rex2: Towards generic object detec- tion via text-visual prompt synergy.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection T-rex2: Towards generic object detec- tion via text-visual prompt synergy

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:11.984434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:11.984434Z digest=sha256:efbe35cadc011245c82854dfe8488dd61c38c1a2a7276291b5ad3ce216851067

Observation 084166b9-93db-4129-8882-125be8f4183a · outbound

This paper cites Prompting visual-language models for efficient video understanding.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Prompting visual-language models for efficient video understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.802754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:11.988563Z digest=sha256:718da74bcd4d7264a0b1db48690e4ccc730a5e94aa73d883d085941dd7655602

Observation 178afd0a-52a8-4b11-b021-c28885dd72f2 · outbound

This paper cites Dense-captioning events in videos.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Dense-captioning events in videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.787985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:11.992614Z digest=sha256:0a3cbb577a79da3a03cefa6241ba4fefb1578e3cdf623629a9f2339d98c1154c

Observation 4ff0dbfe-99f5-4758-a122-9aeb10c15eb2 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection VideoChat: Chat-Centric Video Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:11.996936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:11.996936Z digest=sha256:bf9e399969093b3a3452035b6dfa34d4637edcc13390fc5ce60009115f7abfca

Observation 858e4a86-03dd-4367-9de3-edb0737d1eba · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Unmasked teacher: Towards training-efficient video foundation models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.772567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.002095Z digest=sha256:c0bd8c0dc121408b307f7f478734c3a6b6669c489d56868567f078a92c5a870a

Observation 2a6bc45b-887f-407f-b02c-6b591c815850 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.759155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.006085Z digest=sha256:dd89549be9789c0d734ec7bc9a4a80b29368af82f3037eddc1612801452046cf

Observation dddf181c-cfef-475b-a87c-eae5063a370a · outbound

This paper cites Grounded language-image pre-training.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Grounded language-image pre-training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.745672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.010107Z digest=sha256:5d7258c5efb317a5f497f64ec6792d5da28aebcc5b3064e7947c96bb7d70bea8

Observation e2326f9c-a21e-4f9c-9d64-796a67edefe8 · outbound

This paper cites Detal: open-vocabulary temporal action lo- calization with decoupled networks.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Detal: open-vocabulary temporal action lo- calization with decoupled networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.731870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.014198Z digest=sha256:83441000b91d4eabb42b4b645ab0d8087c527f544f3b8d92b946d3deb401999f

Observation 6a64592b-ffe5-4217-89ae-bc456465e19a · outbound

This paper cites Learning salient boundary feature for anchor- free temporal action localization.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Learning salient boundary feature for anchor- free temporal action localization

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.717853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.019042Z digest=sha256:329f8ec895ddb90d28bebb32249ae5f65a449c51f2cc7d5f5239e30710cd19c9

Observation 357ce09c-b620-473f-bb54-4355dd7f0a32 · outbound

This paper cites Tsm: Temporal shift module for efficient video understanding.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Tsm: Temporal shift module for efficient video understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.023159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.023159Z digest=sha256:f698ea8cf739c687f63ba1a9fa88a0bd3167abebc719075063db541f0776c6df

Observation 1e3ed86d-c28e-45fb-9936-2273393539ee · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Univtg: Towards unified video- language temporal grounding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.027367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.027367Z digest=sha256:20f0cd18f45c99f8f602eccdd9ee696f590a0f715b57146d4a820c0ac5f55a41

Observation 2347ab74-c07c-4f6c-84f6-0ae9820ffe81 · outbound

This paper cites Single shot tempo- ral action detection.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Single shot tempo- ral action detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.686993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.031568Z digest=sha256:d8e91e952f6093b4e298bb9b6580c66401b91ace733b89ab755445317090d6d2

Observation 8656353c-c7b3-4b6d-a36d-4f8a7e93d3c0 · outbound

This paper cites Bmn: Boundary-matching network for temporal action pro- posal generation.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Bmn: Boundary-matching network for temporal action pro- posal generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.035641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.035641Z digest=sha256:f683773dae2f59992163b11bd825c32a538a1324b6338fb8cfb9b8368be84321

Observation 47621ad1-fd99-4c7e-b9f6-bb4c2a3d2bfc · outbound

This paper cites Focal loss for dense object detection.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Focal loss for dense object detection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.039984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.039984Z digest=sha256:85233fe634c6fca47bd5d5f7b3f28254dd43c5e7326b6ed6b18d661d25e95602

Observation bf5fdd36-f2d4-4aff-8841-0fe11f45a718 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.657546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.044191Z digest=sha256:04763bcf3fe29d769a43eb8a11d73e592b314e8843f10d4215577bed2f8daf0b

Observation 3b4eb629-b65c-4e19-ae26-6ee52a125700 · outbound

This paper cites End-to-end temporal action detection with 1b parameters across 1000 frames.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection End-to-end temporal action detection with 1b parameters across 1000 frames

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.644427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.048533Z digest=sha256:b7c76fdd0c8105dfc41b6d0357ccd25b3b4d04a38554ab0cf957b9ed5a3e4b35

Observation cd97fa05-a8d4-49d2-9f87-c17f7b56e8a6 · outbound

This paper cites An empirical study of end-to-end temporal action detection.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection An empirical study of end-to-end temporal action detection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.630764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.052601Z digest=sha256:a55a02dc691c54f20b6a9d725513e734f98e4ecece84a2c4e5a28c63bdcac1d7

Observation 18234106-3f50-4670-81aa-8f358e43d1a9 · outbound

This paper cites End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.617840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.056818Z digest=sha256:f992d3c33625d6d338bd65ebd9955a6ac343c4c446759f8a4bed03da99f77be9

Observation 38ffb9e2-8104-4f0b-9a22-8f88d600f0ff · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.060843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.060843Z digest=sha256:17edfc2a47287ee9e5d57375dff8fa2de0b0ee332fc0c46dc540f908e11e8203

Observation 32f8dd08-93bd-472f-b755-fbb32083ec25 · outbound

This paper cites Fineaction: A fine-grained video dataset for temporal action localization.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Fineaction: A fine-grained video dataset for temporal action localization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.596485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.065214Z digest=sha256:8b93466e14029e02ead7a5d99e400b968c65166a5bf041f907689d3c27b9332b

Observation 39fbbbe0-5869-48c1-96d7-3d761f13c087 · outbound

This paper cites Decoupled Weight Decay Regularization.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Decoupled Weight Decay Regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.069248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.069248Z digest=sha256:12e3da2bd3609af9516a62ddd04198487e6c7faf815136b5e128ce1d000b2d5a

Observation c560b90d-9e59-4f4c-804b-1234c681e8f8 · outbound

This paper cites Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Towards generalisable video moment retrieval: Visual-dynamic injection to image-text pre-training

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.582753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.073330Z digest=sha256:d64f56499272a4a7a1b3c4d0e205f027333900983320610c6d8463c73ab1e659

Observation 2e37c068-59d9-4c0f-9b1d-cfd7baf896a2 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.077492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.077492Z digest=sha256:87ed7a002e9f571d986124ff779bd8747c6a7eac278a9474ff17d8c3f1295f3d

Observation 93bcd896-c203-4d59-ab08-f32a11c4296b · outbound

This paper cites Local- global video-text interactions for temporal grounding.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Local- global video-text interactions for temporal grounding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.082162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.082162Z digest=sha256:3fdbb07c0bb817115a40d0c1c31b63e6112c9590a8a3c25efd9558dab344290c

Observation ccc8222a-f8fc-4507-aea1-9c0035738b10 · outbound

This paper cites Zero-shot temporal action detection via vision-language prompting.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Zero-shot temporal action detection via vision-language prompting

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.559976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.086313Z digest=sha256:f3396f28178c6202c4bef1a7a8554db5c59836cc6fb2e9b9723925a48646d5c4

Observation 53bcc1ca-7d58-498b-8e3e-6180ca02c5f5 · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.090885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.090885Z digest=sha256:3da1eff07954a90825e7437873dfbfe483519c86cfbe9bdb86ef845ea8ef3be2

Observation 704614e2-b4dd-45cc-ab92-4d7fb5de9683 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Learning transferable visual models from natural language supervi- sion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.095814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.095814Z digest=sha256:97a14ed14379d0f17f75f02f6ef01ff9fbe8510625788fca90439e2d24c0bb64

Observation ee6e1fd8-a83c-4c4d-aea8-76e3142a7f0a · outbound

This paper cites Action sensitivity learning for temporal action localization.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Action sensitivity learning for temporal action localization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.099801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.099801Z digest=sha256:3c09472e25c7e6164990b8c66a5ad71ff78960adf2781499267c1b9c6a5c43c8

Observation aca31db7-1f00-475c-9d08-e9018784dd56 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.529824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.104055Z digest=sha256:1dfc02c734dbce2b791ca81a6e43074dea7961a4f922bc8f918761a6031878c9

Observation db9fe65e-e596-4c62-bebd-c5361b85c16d · outbound

This paper cites Tridet: Temporal action detection with relative boundary modeling.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Tridet: Temporal action detection with relative boundary modeling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.516305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.108083Z digest=sha256:08731a1eeb2aa3b6751b2a02438aaa241e4f2cd6a7dddbe53dbc75c615f299a8

Observation 8772812d-fe1e-4bc0-a66a-2b73c5e46dd3 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.112142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.112142Z digest=sha256:da03070e7701cc393673c914216d8f53ac7ea55db0c616c45a4e723a1d943114

Observation 941ccdad-9b86-4eaf-b617-de3919e7f0cc · outbound

This paper cites Attention is all you need.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Attention is all you need

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.116244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.116244Z digest=sha256:9ffffef3a872f488264dd0753ce6b00f09460f3de295a315019cf7986746fe40

Observation 537cd032-6373-483a-825c-aab3c70f68b2 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection ActionCLIP: A New Paradigm for Video Action Recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.120166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.120166Z digest=sha256:4fabacb2ceb7b589377bddfd449a8fe8c2f0864f197ee171e98620a9718550b6

Observation 6d8d9532-581a-4436-92de-8396b86cd79a · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.124563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.124563Z digest=sha256:c2ec85177324984148d36a850f87959ba053247cf4cad058063a25eb3f0e0657

Observation 739c4d95-ae11-437e-ab3d-0315a537f523 · outbound

This paper cites InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.128948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.128948Z digest=sha256:ef0cb75fa53902ff00cbf24b5ff610d75399915d3c3d5382d3b7e87b4849aabd

Observation 9679374f-8457-46e5-a818-bda04611e9a7 · outbound

This paper cites Learning to refactor action and co-occurrence fea- tures for temporal action localization.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Learning to refactor action and co-occurrence fea- tures for temporal action localization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.485818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.133094Z digest=sha256:82b44e439d5aeb6ea207b456ad6c122df5939583d12f7676d59f48b4bfb23cb9

Observation ae7d4550-78fd-452f-b34e-9ac9f0ff334f · outbound

This paper cites Unloc: A unified framework for video localiza- tion tasks.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Unloc: A unified framework for video localiza- tion tasks

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.471428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.136883Z digest=sha256:6b99e97147da73f14162e77cad32a26c536d127e3d4c9c5dd025730d6b1a333b

Observation 900d7373-058c-4c5a-84e0-1a4d3b8d4034 · outbound

This paper cites Basictad: an astounding rgb-only baseline for tem- poral action detection.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Basictad: an astounding rgb-only baseline for tem- poral action detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.457825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.141418Z digest=sha256:6b8caabcde8b123b1abbdb6106f510493587ab3ed9a99667d654e86d330cc609

Observation dbaeb6b5-e803-49a0-bd41-e7ae532faca2 · outbound

This paper cites Detclipv3: To- wards versatile generative open-vocabulary object detection.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Detclipv3: To- wards versatile generative open-vocabulary object detection

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.443421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.145534Z digest=sha256:dffc3257ff123c996a4189380cc314b8a7e5c551bb8f56968eea5e3ecca57811

Observation fe123d0a-b629-4d54-9574-a435d9734cd6 · outbound

This paper cites Semantic conditioned dynamic modulation for tempo- ral sentence grounding in videos.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Semantic conditioned dynamic modulation for tempo- ral sentence grounding in videos

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.150258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.150258Z digest=sha256:8c35dd5cf4016b02db2eef002dffd6291aa3d455e355610609183c055f3fb159

Observation 8bfbcb11-769e-4293-8ec2-9b0570a481b5 · outbound

This paper cites Graph con- volutional networks for temporal action localization.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Graph con- volutional networks for temporal action localization

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.421012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.154906Z digest=sha256:cec90eab45a6aab4de3dd223327892cb626ea0e498ca833566de34b1b258c9ef

Observation e8639064-53de-4d74-991a-83d9d2e4c7cb · outbound

This paper cites Dense regression network for video grounding.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Dense regression network for video grounding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.159390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.159390Z digest=sha256:6895986c32bd0ecb226dbccdbac21c0e5026cf2759bc2b8678218303e9398bc6

Observation df438b76-7cb8-466f-9a6c-b5bc4ccfbe76 · outbound

This paper cites Unimd: Towards unifying moment retrieval and temporal ac- tion detection.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Unimd: Towards unifying moment retrieval and temporal ac- tion detection

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.398852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.163620Z digest=sha256:e9cabbe14dbfcfd6dbf39c5bc954e9275867750d7c2e814e04b943109e02ad5f

Observation 982e96c6-3a4d-47e7-b8f3-56993819cb5b · outbound

This paper cites Actionformer: Lo- calizing moments of actions with transformers.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Actionformer: Lo- calizing moments of actions with transformers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.167717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.167717Z digest=sha256:e60add90bd53cef04daa9a7038337439ef51c62bfb174ecc2902132f73516ed1

Observation a5c32f57-678f-4876-9b73-922107083c3f · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:50:12.171758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:50:12.171758Z digest=sha256:b6e8081abb246cff27cf4ba6ba90f69515ffa195fe4af761206347f0f8d96e90

Observation 03f4f722-528f-4aef-af00-43f5aabe29fa · outbound

This paper cites Learning 2d temporal adjacent networks for moment local- ization with natural language.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Learning 2d temporal adjacent networks for moment local- ization with natural language

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.376418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.175940Z digest=sha256:54852be5c580362b3d9bb91581d10f668f103e5f00797fea8bfaee0d64a5e173

Observation 60f187a1-fe2d-4004-abd5-18bdc29a4533 · outbound

This paper cites Hacs: Human action clips and segments dataset for recognition and temporal localization.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Hacs: Human action clips and segments dataset for recognition and temporal localization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.362814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.180563Z digest=sha256:c5c4b645654de92933ac1989ee2adbefd20962a81eff1ad362077bdfacd0a394

Observation d25f12e9-52ae-4d3b-b006-57dde7b74ee4 · outbound

This paper cites Distance-iou loss: Faster and bet- ter learning for bounding box regression.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Distance-iou loss: Faster and bet- ter learning for bounding box regression

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.349016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.184512Z digest=sha256:9a387b7f8fe89f25613dc412f74baf21c3bf8cf59085ee31c0da0d550154555a

Observation a39519a1-9e4f-4c62-8deb-c1b9232acb08 · outbound

This paper cites Enriching local and global contexts for temporal action localization.

Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection Enriching local and global contexts for temporal action localization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:50:12.334504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T11:50:12.188418Z digest=sha256:d80429c40805436a4ab8e295aaceb061b21fb22fcf6f5710c028dcbf25159f8a

Pith citing papers

No inbound Pith citation observations are available.