Pith. sign in

Paper Citation Record · LEDGER

Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2109.04290.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.04290 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:41:46.455487Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T14:27:02.995751Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b57751d6-c40e-4a3b-b15e-d4494cb599e1 · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.705835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:d6ab1356828e416bc1a9d71af68528a8cfcf63092281851388e77c74146fe09c

Observation 779a5061-293a-4889-b7ae-50a2d3404ad7 · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.304132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:7551b61ec6e7c22a8f5568e3c3306c22846a415c778aeaa155bf59cfc4cd593f

Observation 705ea039-16c3-467c-a9a3-62d22b2d774d · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.302325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:e97f1f5289f145e45e923ab0b98e4278a8a58d7dea98675e16a90c3ab12c09de

Observation 87e83368-c242-4fc3-bae2-17c0f6ba70ab · inbound

Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review cites this paper.

Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:46.455487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:46.455487Z digest=sha256:227c04e803f61111e714d79b84b91f0489d7d3d24f62fc8e0d861a75cb3a91b6

Observation 32d8a49a-cd82-4d9a-ac3e-fd9236325f46 · inbound

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval cites this paper.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.521872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.521872Z digest=sha256:f467bbf771f9e9967138dbf99b331c519590c4ea08a8251a58468523230fb875

Observation 45da16d9-30be-405e-9eac-6718a94c4c42 · inbound

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis cites this paper.

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:06:09.617931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T15:05:37.964883Z digest=sha256:188d456c3f77c336590c5a7d62efa58bbe97df14d470035aedaee1e0ae0f9c50

Observation b28d9ea2-d15a-49ba-b184-ebb6c5bfb2c2 · inbound

MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment cites this paper.

MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:27:02.997566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T14:25:04.811472Z digest=sha256:9bc156590c4b19182d4835f70467dcedf0ea9e76ce9b2553a830a36723dbce4e

Observation ee3dbced-33bf-42f2-9c68-e71ac961bc71 · inbound

PHA-Net: Prototype-based Hierarchical Alignment Network for Text-Video Retrieval cites this paper.

PHA-Net: Prototype-based Hierarchical Alignment Network for Text-Video Retrieval Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T00:46:15.864440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:46:15.864440Z digest=sha256:788c5d8107caef6f5dca34cc5df9ec3b616c4b5e59de8684a08311168af7979d