Pith. sign in

Paper Citation Record · LEDGER

Span-based Localizing Network for Natural Language Video Localization

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2004.13931.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2004.13931 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:13.924824Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:14.529226Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation efadc4be-a648-4ab5-bd4d-166249e945d8 · inbound

DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding cites this paper.

DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding Span-based Localizing Network for Natural Language Video Localization

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:53:54.016914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-24T04:51:07.439654Z digest=sha256:ef8a547aa01b552381c121164404966a612fbc5786e676ede60cab1fe8e1514d

Observation 15d63241-d3b1-4a46-a0c3-6eb7bfb71829 · inbound

Multi-Scale Contrastive Learning for Video Temporal Grounding cites this paper.

Multi-Scale Contrastive Learning for Video Temporal Grounding Span-based Localizing Network for Natural Language Video Localization

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:32:43.053634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T07:29:46.573683Z digest=sha256:d9d88e25a9283f75c8b449db9efb9853dbca6908a37d1cbd3944afbbe3577302

Observation 1f2ab522-d356-44f2-9857-32b249d58841 · inbound

DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos cites this paper.

DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos Span-based Localizing Network for Natural Language Video Localization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:13.924824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:13.924824Z digest=sha256:9cfd05a82cfed7448edfb1fc63a77e820fbced6daf8c2456e9668e27063a34e6

Observation 8fa31332-25ce-4b36-930f-88143aa3d190 · inbound

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering cites this paper.

MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering Span-based Localizing Network for Natural Language Video Localization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:29:28.693277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:29:28.693277Z digest=sha256:40305fd197ae99651767b19c2171cbcca0bd9ee1334503ffbdbed2a36c697b18

Observation 69d7761d-49f7-4f9d-bb83-9b7e0e341bfc · inbound

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning cites this paper.

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning Span-based Localizing Network for Natural Language Video Localization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T17:01:56.789076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:01:56.789076Z digest=sha256:11e737ebb3c73c0c196e141b0101470bd0fa1e412bfee54ae7345bca892f774d

Observation 8d2e8ea1-3eb3-4cce-b704-39e7b0c783fe · inbound

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting cites this paper.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Span-based Localizing Network for Natural Language Video Localization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:58.184421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:58.184421Z digest=sha256:b44ac7293beb5d67bfc566cd6eb4eca223e5d18d8063621cd208b8fb2b764898

Observation 5dc1d1af-fa7d-466a-a405-990cb0ba80c6 · inbound

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding cites this paper.

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding Span-based Localizing Network for Natural Language Video Localization

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:51:02.385460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:55:56.385801Z digest=sha256:aa26ae7941913ea3c2321b16d1caa5d58e8e070b83166e64290bc7510d287a17

Observation da339d9c-afb7-4394-ad2a-e7de88a1e12b · inbound

MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding cites this paper.

MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding Span-based Localizing Network for Natural Language Video Localization

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:51:31.166366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T01:30:17.546118Z digest=sha256:febdd79e182eacd5b0ca8a218a265f33357f9dc50bcc862cd70b07b68647bb72

Observation 633bdd30-49df-46b2-bf1a-5b28286d07f1 · inbound

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding cites this paper.

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding Span-based Localizing Network for Natural Language Video Localization

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:32:52.529810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T19:29:47.356665Z digest=sha256:8b61a9eb0f092cdba45f3b9ea2aff677b69b8c9cb9729e6c6d90325abbd7fc67

Observation 260a7219-af8f-4735-87bc-eae83eed1ee5 · inbound

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection cites this paper.

CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection Span-based Localizing Network for Natural Language Video Localization

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:14.530960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T17:34:40.317688Z digest=sha256:bd07785b02a0957c8783935e24541c94bd6a989e9f197d4292a981ce7c9c0b5b