Pith. sign in

Paper Citation Record · LEDGER

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection

As of 23 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2606.01149.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.01149 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T17:34:40.317688Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22a30a15-8677-4417-9d42-539516125641 · outbound

This paper cites Learning 2d tempo- ral adjacent networks for moment localization with natural language.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Learning 2d tempo- ral adjacent networks for moment localization with natural language

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:48f1dca72a1f7ab7210e295e9ed588df6d4ec780c23e6d0e211a58d42e1298a5

Observation 315786f7-fc8c-4ef3-89f7-476f7a60efb8 · outbound

This paper cites Cross-modal moment localization in videos.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Cross-modal moment localization in videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:58f8a20d6737ce08f3a269f48ab1af9e72b42e8949d523a6e2f70b985941fd21

Observation 93bdc89b-1ed2-43b1-87b3-3f65931dbf00 · outbound

This paper cites Tall: Temporal activity localization via language query.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Tall: Temporal activity localization via language query

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:d8fe2494b380af10cea9c13d68d2984759f7702cad77786caafa16e638b5e617

Observation 807a686c-b33a-40e1-8932-7b1173abde40 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Ego4d: Around the world in 3,000 hours of egocentric video

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:ca6b892ff1159dcc4df1b583d1d39e6ba53b8a4cb3854f84c6a9ef313da0dfae

Observation b4e9b3c8-39ed-463b-8411-a0905c6a0e82 · outbound

This paper cites Less is more: Learning highlight detection from video duration.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Less is more: Learning highlight detection from video duration

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:06de967cabb6f29dabf65ad7103c110e26e21339780aaafd06781d3bcb97e614

Observation 89c41202-bb65-47c8-a209-7a234c1e3025 · outbound

This paper cites Highlight detection with pairwise deep ranking for first-person video summarization.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Highlight detection with pairwise deep ranking for first-person video summarization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:c3b71b023d9e06f053131af227f6990b082b5f5e55114a0d5bab9e4e5901daaa

Observation 3640431a-88b3-4b14-8cf3-e0e5d4455b8a · outbound

This paper cites Localization-aware multi-scale representation learning for repetitive action counting.VCIP, pages 1–5, 2024.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Localization-aware multi-scale representation learning for repetitive action counting.VCIP, pages 1–5, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:e536b51d8950f70637c62254a7d3d3027169ad76966dd86fa8e3905313c52f4e

Observation b014f268-1107-4027-a5de-12badeaa476d · outbound

This paper cites Lavt: Language-aware vision transformer for referring image seg- mentation.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Lavt: Language-aware vision transformer for referring image seg- mentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:e97e12eb3d8d3286337982728721d51cfd2ebae09e298f78276b15ac33802a0e

Observation c63caecc-aab3-4a68-bef0-03d1df34abbd · outbound

This paper cites Univtg: Towards unified video-language temporal grounding.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Univtg: Towards unified video-language temporal grounding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:87f79b9da8045176d7b6d373fb83144bc6a4fdb45770f4466c69e77911887c17

Observation e1ae82c3-edbe-4f75-97c1-deea0dc54893 · outbound

This paper cites Detecting moments and highlights in videos via natural language queries.NeurIPS, 34:11846–11858, 2021.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Detecting moments and highlights in videos via natural language queries.NeurIPS, 34:11846–11858, 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:45e33f5781503513a05930ca8147a1e364544cc61e47c860adccb411b6a778a3

Observation 569c3a6e-ca37-4806-bab2-28ef02604c43 · outbound

This paper cites Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:14.539786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:a8d9e4c5151a1d30615948a797ad76c907051e0b34254cd5795b5414414e6f0b

Observation dcfba0e6-3f5a-42c2-a300-70b41dd16d48 · outbound

This paper cites Bridging the gap: A unified video comprehension framework for moment retrieval and highlight detection.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Bridging the gap: A unified video comprehension framework for moment retrieval and highlight detection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:c24364d200c2f62061d2ce3ef569e114d30c1bc8eb242acbbb32495a0ea1f999

Observation 2014460d-7925-4b21-aaf6-760aa3a13b00 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:ccbe9fe89cc7c2d328e8671c4501a6eb0916c1118c2373e9c1fb5b8ec48c07f3

Observation eede6fcc-7381-4bb2-8dcc-18519b08daef · outbound

This paper cites Cva: Context-aware video-text alignment for video temporal grounding.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Cva: Context-aware video-text alignment for video temporal grounding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:ece480ffe8a50e8cac72f7193567bdf4c552ae74ba85e0545bcaa2bcf9ffd6c1

Observation 7b974dcb-e8a3-42ac-a021-46aeb014980a · outbound

This paper cites Timeexpert: an expert-guided video llm for video temporal grounding.ICCV, pages 24286– 24296, 2025.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Timeexpert: an expert-guided video llm for video temporal grounding.ICCV, pages 24286– 24296, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:2d9a3b3c091d6cc9b0544551ea2057ad430e41544fcdc4cd6a97c5491d50f98f

Observation 6d2bbe2b-bd0c-4857-955f-bf533b06a53c · outbound

This paper cites Multi-stage aggregated transformer network for temporal language localization in videos.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Multi-stage aggregated transformer network for temporal language localization in videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:bed5d5b5907693f2508f661d7248e262c746f3428252b60545f0faba380b8287

Observation c669091b-0c0d-45d0-b96f-ad8c76193dd8 · outbound

This paper cites Semantic condi- tioned dynamic modulation for temporal sentence grounding in videos.N e u r l, 32, 2019.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Semantic condi- tioned dynamic modulation for temporal sentence grounding in videos.N e u r l, 32, 2019

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:0922bbc541a21aed86b47546e935b63713cc702cffb7ea7e5ccad3d936d1731d

Observation ce2713e0-2d7a-4794-b9fa-defbfc55418b · outbound

This paper cites Structured multi- level interaction network for video moment localization via language query.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Structured multi- level interaction network for video moment localization via language query

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:3080f266a03988573dae16f2a9b833bca350a86a1bb0c4a2379428c74773d18f

Observation b239ad8f-e068-49de-9c36-33222c410988 · outbound

This paper cites Video2gif: Automatic generation of animated gifs from video.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Video2gif: Automatic generation of animated gifs from video

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:d50217d08555373019a32ec4743a61a611de57b60bcf0069d791eec7ecce5840

Observation 17b6f008-334b-4cd0-a096-e96c48a734ab · outbound

This paper cites Video highlight detection via deep ranking modeling.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Video highlight detection via deep ranking modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:0a42cf6e15a6e4b1ca3fe0ea68e7a2438cac711e50d1d5f5b45602136aa70e08

Observation fe1d3656-4b88-4bf1-9fa2-209011b34fa7 · outbound

This paper cites Joint visual and audio learning for video highlight detection.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Joint visual and audio learning for video highlight detection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:e0e1f4a1d80d143d4fd948340687dbb9b8a383e00d4782a7ccd0a164030b62a0

Observation a84a6e54-d1b9-4caa-8ef1-e11cfc541210 · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:b003db7202e2617878273652375de2c0b4f7113f83a37e8beea09013f569ff7d

Observation c2b60d9b-d6ab-47c5-9dd1-32c47faf604e · outbound

This paper cites Attentive moment retrieval in videos.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Attentive moment retrieval in videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:5755bf700ba0a27e7068074ce15bcc59276686cfc3792b142bce5d18c8cf7959

Observation 682b1ace-b821-4715-a363-8a1ab5d6f8c5 · outbound

This paper cites Localizing Moments in Video with Temporal Language.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Localizing Moments in Video with Temporal Language

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-01T20:56:14.537192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:4bff55eb049f2a1b335e49f1490fa5b6b59ad02969e54b7201cb62b0a491647c

Observation fd65be01-a2c6-4484-9bde-7e789cdaa36e · outbound

This paper cites Proposal-free video grounding with contextual pyramid network.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Proposal-free video grounding with contextual pyramid network

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:b5f9e792fce284fcd74ad338375465b96fb46a6e6b59881782e5ebd4615f6f46

Observation 019d4611-d1c9-46e7-8327-5b92d9929a4b · outbound

This paper cites Local-global video-text interac- tions for temporal grounding.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Local-global video-text interac- tions for temporal grounding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:45e72209be2aa2e7b388a17dc8c09d3818ed062fbecdfb21629ddeb25df4112f

Observation 50d7b8a5-a007-4b1f-a49c-71bbd569c332 · outbound

This paper cites Proposal-free temporal moment localization of a natural- language query in video using guided attention.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Proposal-free temporal moment localization of a natural- language query in video using guided attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:0a189f856b1016dd93f9204c2984c0bc73f8ddfbd76c8f6f5096baaee56579be

Observation 03e962f4-8e62-4924-a5b7-ae318afa68e3 · outbound

This paper cites Frame-wise cross-modal matching for video moment retrieval.IEEE Transactions on Multi- media, 24:1338–1349, 2021.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Frame-wise cross-modal matching for video moment retrieval.IEEE Transactions on Multi- media, 24:1338–1349, 2021

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:65d45860202cf6f23bb11d6f7f23398e478ce76c56ab501447b61070517d0ba7

Observation baddb296-b29d-45a8-a6bc-f511dab6999a · outbound

This paper cites To find where you talk: Temporal sentence localization in video with attention based location regression.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection To find where you talk: Temporal sentence localization in video with attention based location regression

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:210ce382847d4f5a526a51e4be0215cd92cbb29dfe90b260cec6b6f17ebd6655

Observation 260a7219-af8f-4735-87bc-eae83eed1ee5 · outbound

This paper cites Span-based Localizing Network for Natural Language Video Localization.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Span-based Localizing Network for Natural Language Video Localization

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:14.530960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:c9feb9d19c6b59c2539337f1b174b4fb1eb7cd1579786120a7aab7129a36ff6c

Observation c27b6fea-5d48-4c75-b3de-52c86e2f92d9 · outbound

This paper cites Ranking domain-specific highlights by analyzing edited videos.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Ranking domain-specific highlights by analyzing edited videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:44f7c0b30dd4c7a8cbe28c0343a27c0b9b0659cd6bd3396fe81e35e0cb997b91

Observation 95fc8d5d-7119-4369-a1f5-a4e774b2df86 · outbound

This paper cites Tvsum: Sum- marizing web videos using titles.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Tvsum: Sum- marizing web videos using titles

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:5d33c0f57c1368e00dd6febe795e8e21f5db3cdde21ca998dee348c4837526e9

Observation a257796b-3d46-45c3-9f2d-817c280f2792 · outbound

This paper cites A deep ranking model for spatio-temporal highlight detection from a 360◦video.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection A deep ranking model for spatio-temporal highlight detection from a 360◦video

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:2067424f25e627ab4e8f03f969c2950e02841147b647bba22856425d06df9cf7

Observation 224309cf-1e62-4668-be3e-141ff638de25 · outbound

This paper cites $R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection $R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:14.534958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:fdbe343387423135b5cf8044692fed948b931201601879c489222fc17bf00eac

Observation 3f848c19-5194-4a95-b227-c800c35b7bb1 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-01T20:56:14.533265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:25b098825fde0047f7ad3c838ca7462a408cb810909ff8d07aecfe43dd1ad27d

Observation 872b2b73-8484-4820-a8e3-aacb961627a5 · outbound

This paper cites Grounding action descriptions in videos.Trans- actions of the Association for Computational Linguistics, 1:25–36, 2013.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Grounding action descriptions in videos.Trans- actions of the Association for Computational Linguistics, 1:25–36, 2013

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:4071e473b165fa11a37ec2e9479a4da15517f1b53596aa5edf2ab80c36c6edc1

Observation 3630ce14-89ba-4ba9-b58c-56b0d6afbdb8 · outbound

This paper cites Mh-detr: Video moment and highlight detection with cross-modal transformer.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Mh-detr: Video moment and highlight detection with cross-modal transformer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:910f1cc30f03d2311cc4dc7a2038e65c40671e66fa6769534011d7f99e7bf12d

Observation 8ebe4664-1556-4ed5-b25b-276834e71b3c · outbound

This paper cites Knowing where to focus: Event-aware transformer for video grounding.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Knowing where to focus: Event-aware transformer for video grounding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:1dd844dfbde18992b7ce8e07387f25d048e7767d84942410234dd07e5c7ce6b4

Observation b7aa1ef9-9934-4632-8560-75c88061515b · outbound

This paper cites Query-dependent video representation for moment retrieval and highlight detec- tion.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Query-dependent video representation for moment retrieval and highlight detec- tion

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:49740b78b09c947b348a260c81d32cf7c9c6bc1371687a611549cc65e511716a

Observation 58d65551-7271-4859-ae98-a7305eac83f4 · outbound

This paper cites Video summarization with long short-term memory.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Video summarization with long short-term memory

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:49f69648c72c12846856d5964f1bdc186cd1ec1db2aa8079cb98e33bcd345616

Observation 9e9fa9b8-03df-4717-a239-fe8b23ba8617 · outbound

This paper cites Learning trailer mo- ments in full-length movies with co-contrastive attention.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Learning trailer mo- ments in full-length movies with co-contrastive attention

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:9c0463d89ceb2ba25c76d7e3444d128c543a96b63fc37e08e475fc782501e2ce

Observation b9b02cda-98be-41e7-80ac-728539a6f863 · outbound

This paper cites Cross-category video highlight detection via set-based learning.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Cross-category video highlight detection via set-based learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:85f8158044ca06ecd53920f85edc126943387f2e2699b9f0bb26b77840be8498

Observation 7bbc33d7-f84e-425c-9821-9343037ff7d5 · outbound

This paper cites Mini-net: Multiple instance ranking network for video highlight detection.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Mini-net: Multiple instance ranking network for video highlight detection

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:fa84415bbdbce3de4a79ce32d2171f8c40dff8f7d483f2bca41c855ebe1f3a9d

Observation a6e022f1-8fba-438d-a54b-74d5ab28a1c2 · outbound

This paper cites Temporal cue guided video highlight detection with low-rank audio-visual fusion.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Temporal cue guided video highlight detection with low-rank audio-visual fusion

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T17:34:40.317688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:bbc5afa0f8a17ecf0688def369a2b442eee5720a00cd1cc73fbbef0a4b4adbc7

Pith citing papers

No inbound Pith citation observations are available.