Pith. sign in

Paper Citation Record · LEDGER

VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.00741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.00741 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:29:59.081986Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:09:44.533509Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9cfa4a79-8cc5-4f7c-b69e-5a984caa60a7 · inbound

VIRES: Video Instance Repainting via Sketch and Text Guided Generation cites this paper.

VIRES: Video Instance Repainting via Sketch and Text Guided Generation VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:29:59.081986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:29:59.081986Z digest=sha256:73df3c76ff257489f377cd65d6ff897bc3e9da4dd05609ee0b4e415f12bbd5e5

Observation e8bd4cc6-39d3-4cbd-8c24-ed4b64bf2f5f · inbound

PanoWan: Lifting Diffusion Video Generation Models to 360{\deg} with Latitude/Longitude-aware Mechanisms cites this paper.

PanoWan: Lifting Diffusion Video Generation Models to 360{\deg} with Latitude/Longitude-aware Mechanisms VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:49.485512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:49.485512Z digest=sha256:ef0f335be20e19c41ccf0ed40bbd02b3756a3e6a64af0cae54d2ee280f3be2d0

Observation 147b979d-ab45-4515-b8eb-aa7a2ccea265 · inbound

ContentV: Efficient Training of Video Generation Models with Limited Compute cites this paper.

ContentV: Efficient Training of Video Generation Models with Limited Compute VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:34.699982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:34.699982Z digest=sha256:e5891c848d6997f4a77002567ed57c3d857bc07dd54c9edee90448c036889abb

Observation ccb49950-47e0-4796-a667-e079589500d6 · inbound

Audio-Sync Video Generation with Multi-Stream Temporal Control cites this paper.

Audio-Sync Video Generation with Multi-Stream Temporal Control VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.577220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.577220Z digest=sha256:1cf3cf374295910e37ee86621da10bfbd0371841e693c13c17cba6979dc9dc3a

Observation 7e6f7a3f-13d5-4433-b15e-9a32d1ba58a9 · inbound

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs cites this paper.

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T20:23:06.982542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:23:06.982542Z digest=sha256:d8157a83cb4690ff9e9305ef117450fcb513f905ef7449c624064fab654dfd92

Observation 25e440a3-5a80-44fd-808c-d66696d922a1 · inbound

AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner cites this paper.

AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:28:40.814334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T23:26:21.855485Z digest=sha256:b9f8e52199de7211b47838f348c0a83cd140f18031b9a38ab6b2e29f41515a97

Observation c57f45b7-3a29-4280-b431-449a958ab4c1 · inbound

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation cites this paper.

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:38.543505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T18:21:00.089755Z digest=sha256:8ae28ea39b924f0dd762d54188fc9b865404e2bf82d930a09cda8a512d16121a

Observation 1d00a4cc-e020-4ef3-af36-f28fafd198d0 · inbound

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries cites this paper.

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:26:25.331526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T05:22:14.870351Z digest=sha256:7a1f3697f0d4edc1f932dada99023b7b78337aad457af3fb66aaa91cd15b7547

Observation 32a1a612-157a-42bc-b682-af9de09616ef · inbound

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars cites this paper.

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:44.534990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T09:09:06.925645Z digest=sha256:7fe2957be4395f83af0796bc9d792fa350543a0ec58aeed9454b0c0e7244ee23

Observation 8643a6bb-73e4-4e36-8814-f93bf3e22a6c · inbound

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars cites this paper.

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:05:29.176076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T07:00:53.496569Z digest=sha256:7175507d8c167c1498ab1a1147b9f298c0ed433ec071f3d5efc51b08056fd0f5

Observation 6b1b4e89-7b7d-4b19-bda2-02134832aaf7 · inbound

InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving cites this paper.

InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:55:40.670808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T06:05:17.158363Z digest=sha256:8868553c7347ba566d301ed0d19d754afdbd04180f1bb2d7c5e797ec1a3c9b84

Observation 85f421ea-4091-4a7b-82a6-397c160dea31 · inbound

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval cites this paper.

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T18:30:14.584902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:30:14.584902Z digest=sha256:451b7302ee8f12d20967e64bbc9c351f3753ec41c7188f0fe211fd795764d3fb