Pith. sign in

Paper Citation Record · LEDGER

Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2109.04290.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.04290 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:48:37.065489Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T14:27:02.995751Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b57751d6-c40e-4a3b-b15e-d4494cb599e1 · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:00.705835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:79f1c85a1a9d67770129002ad58e48dbea82b877d912ec30253d8f006eabd19b

Observation 779a5061-293a-4889-b7ae-50a2d3404ad7 · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.304132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:e709b842b403efb2f9c7bcd746f707de6371eb5acfc8f9ba0d996854221d7aaf

Observation 2e0d8cbc-c51c-4c2d-9c74-022e3fbfc770 · inbound

GIMS: Image Matching System Based on Adaptive Graph Construction and Graph Neural Network cites this paper.

GIMS: Image Matching System Based on Adaptive Graph Construction and Graph Neural Network Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:01:19.543892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:01:19.543892Z digest=sha256:1990f146b0942eb3742ed4cfe42c3a7cc2bf71ba266309322ade1d8138ec8de7

Observation dd9b1ac8-f38d-49ec-ace9-3cf8cf4551c1 · inbound

Language-based Audio Retrieval with Co-Attention Networks cites this paper.

Language-based Audio Retrieval with Co-Attention Networks Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:11:53.248081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:11:53.248081Z digest=sha256:507299b34e1255b5314765337cc809842fc5d092e1b8b2d6473ee78746c668ad

Observation 705ea039-16c3-467c-a9a3-62d22b2d774d · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.302325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:ba6aec470d29925e7952a16fbba52faa042e07822cda48a85d31a8b434aca52f

Observation 87e83368-c242-4fc3-bae2-17c0f6ba70ab · inbound

Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review cites this paper.

Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:46.455487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:46.455487Z digest=sha256:7d58348ba0081fbe91d4a0fcc30fb736498a872b6d698c49183dd8588b1caf93

Observation 32d8a49a-cd82-4d9a-ac3e-fd9236325f46 · inbound

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval cites this paper.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.521872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.521872Z digest=sha256:d08e8838c8559be65ae47bab4364eb52e3c8f3844696d91972b39a505f30fde0

Observation 5cba8cd9-fc80-4d55-91ae-ae9fca97ddf4 · inbound

T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval cites this paper.

T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video Retrieval Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:48:37.065489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:48:37.065489Z digest=sha256:1330f32c681f0d6f26f8f72452c7a79d47504689cf04816830a249fb83b1d95d

Observation 45da16d9-30be-405e-9eac-6718a94c4c42 · inbound

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis cites this paper.

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:06:09.617931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T15:05:37.964883Z digest=sha256:b52d771b65b451c206202f66d2770a9bd29ef6768a7091ee09e879a8112cb7f7

Observation b28d9ea2-d15a-49ba-b184-ebb6c5bfb2c2 · inbound

MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment cites this paper.

MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:27:02.997566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-02T14:25:04.811472Z digest=sha256:d757b269b4c462c930479face2078a953c15839c8bfb9901b0be33864f424f57

Observation ee3dbced-33bf-42f2-9c68-e71ac961bc71 · inbound

PHA-Net: Prototype-based Hierarchical Alignment Network for Text-Video Retrieval cites this paper.

PHA-Net: Prototype-based Hierarchical Alignment Network for Text-Video Retrieval Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T00:46:15.864440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:46:15.864440Z digest=sha256:6f3250ddb34911279c0bec67f87930dd61d9c4c886c189c11360c522fb004d85