Pith. sign in

Paper Citation Record · LEDGER

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval

As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2506.08887.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08887 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:04:38.641568Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact1
  • verified fuzzy48
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69266413-019c-46c6-9327-ef3c60df9862 · outbound

This paper cites Localizing mo- ments in video with natural language.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Localizing mo- ments in video with natural language

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.101686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.504246Z digest=sha256:efebb0d774581bafbfd628b78c946f788b53e63f404ebb2d675eca943aeb9661

Observation 8819666a-0375-46b3-b42a-42f9bb34c443 · outbound

This paper cites Vqa: Visual question answering.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Vqa: Visual question answering

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.095804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.507071Z digest=sha256:99146b7745bc3a1aa953aa48d0c7091094f9d7cc2e01eb61cdb51839e1a179d6

Observation 4628a9ca-b838-4b54-b7ce-56bd9964f1c3 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.089994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.509532Z digest=sha256:08980d3cd4d0fb4745b45b0e15e090cdae0235359bc39b84fdcb70f2e6b9ce4e

Observation 0b308772-8625-4b7a-b330-03ca9c3b1748 · outbound

This paper cites Cross modal retrieval with querybank normalisation.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Cross modal retrieval with querybank normalisation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.083821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.512142Z digest=sha256:b60baabb87f64b6101c3887eee9f473bfcac62e6135ebd82b282adb7ace787f7

Observation dc198174-2b83-4f14-a585-9fe2ece5a579 · outbound

This paper cites RAP: Efficient text-video retrieval with sparse-and- correlated adapter.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval RAP: Efficient text-video retrieval with sparse-and- correlated adapter

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.077014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.514579Z digest=sha256:1407a99708bde979a79792ef372d79ba376abd5676f80ebd25468dfcccefb34e

Observation b5ee4d76-0774-4ee6-9b80-77e60927dd13 · outbound

This paper cites Adaptformer: Adapt- ing vision transformers for scalable visual recognition.Ad- vances in Neural Information Processing Systems, 2022.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Adaptformer: Adapt- ing vision transformers for scalable visual recognition.Ad- vances in Neural Information Processing Systems, 2022

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.070737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.516855Z digest=sha256:2a01360df310df173ee315b737c4201be177c8f0c25652e32b9195244ca5cb45

Observation 23c14efb-88db-425d-b8a0-770b3f7f91d7 · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.Advances in Neural Information Processing Sys- tems, 2023.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.Advances in Neural Information Processing Sys- tems, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.064336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.519656Z digest=sha256:94e49c9c3eb05fc30add1516acca3066e22ea431e80a12f7ab1eaef9efedecba

Observation 32d8a49a-cd82-4d9a-ac3e-fd9236325f46 · outbound

This paper cites Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.521872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.521872Z digest=sha256:f467bbf771f9e9967138dbf99b331c519590c4ea08a8251a58468523230fb875

Observation f48e58cb-3cbe-4924-ba15-f4b2fe443d8c · outbound

This paper cites Prompt switch: Efficient clip adaptation for text-video re- trieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Prompt switch: Efficient clip adaptation for text-video re- trieval

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.058173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.524699Z digest=sha256:509f01ca291b4bbc29eaa39dd1f335e95a1c20dc6f0746accafd0be4814d9385

Observation 97dfb5ab-823f-42b0-a475-1ea517437e7a · outbound

This paper cites Acnet: Strengthening the kernel skeletons for powerful cnn via asymmetric convolution blocks.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Acnet: Strengthening the kernel skeletons for powerful cnn via asymmetric convolution blocks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.051684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.526766Z digest=sha256:0ee7696a5cc3f009d2bd6bac9646b2b10c63da355cbf5106c10f4939cd56f0d9

Observation 41a91c59-7797-47da-890b-a138a5bbbdcf · outbound

This paper cites Repvgg: Making vgg-style convnets great again.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Repvgg: Making vgg-style convnets great again

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.044688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.529008Z digest=sha256:c82533ccceee9bc3b85b03e67a093c52cb820e84faf3dbee1932c305fa23d986

Observation 1a91173e-b2c3-4f50-a64c-20c3fbbf3af1 · outbound

This paper cites Improving clip training with language rewrites.Advances in Neural Information Processing Sys- tems, 2024.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Improving clip training with language rewrites.Advances in Neural Information Processing Sys- tems, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.038445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.531373Z digest=sha256:d12a9673edab1ff4ec56e252c6438609d1a2f90507d2eddf99219cdf542106bb

Observation 32c04091-3f0f-4b2c-8cd5-ef0c393c6c98 · outbound

This paper cites Multi-modal transformer for video retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Multi-modal transformer for video retrieval

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.031867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.533670Z digest=sha256:612969414e40eee78071ff4e25c064cc7cccf62bc4d58f6d8793e18151fd8ed8

Observation 9f8a736e-eac1-4dc4-acbb-27e3224e3cfe · outbound

This paper cites X-pool: Cross-modal language-video attention for text- video retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval X-pool: Cross-modal language-video attention for text- video retrieval

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.025277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.535876Z digest=sha256:f14415a31c006cc343138e9edcc8d65dccc6917bb472ee8171fd6e7dc2a178b5

Observation 093d98e5-9f5f-4bc0-af90-1ad33cb684ec · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.538033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.538033Z digest=sha256:5cc296a0be60322cd0a55ae416c747229e5c693a9eac6dcdf22c8e18db59225d

Observation 55a5bf61-fbbc-45a9-8b9b-848a786f3dde · outbound

This paper cites Framewise phoneme classification with bidirectional lstm and other neural net- work architectures.Neural Networks, 2005.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Framewise phoneme classification with bidirectional lstm and other neural net- work architectures.Neural Networks, 2005

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.018752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.540500Z digest=sha256:e0d26cd289c61b2e5d57f994e78a5aed14b6f94cff3f0e3255785b36f1b932a4

Observation 57571e0b-2ffa-47b9-996f-bb3704489bd4 · outbound

This paper cites Towards a unified view of parameter-efficient transfer learning.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Towards a unified view of parameter-efficient transfer learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.542644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.542644Z digest=sha256:4a97e293d2c1e8f37083d80b3c95a52d403a1efc69039decfc5121005e6d301f

Observation 63af92d2-9708-40f8-bd12-91613cb213f3 · outbound

This paper cites Secret: Self-consistent pseudo label refinement for unsupervised domain adaptive person re-identification.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Secret: Self-consistent pseudo label refinement for unsupervised domain adaptive person re-identification

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.008141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.544754Z digest=sha256:daf7ed3959563a505717b4b2e0a3bf88e3bea1e63b8d30a58985cf8827069e77

Observation 3e4554c0-31bb-48be-b182-ab87fffcf3f1 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Gaussian Error Linear Units (GELUs)

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.546950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.546950Z digest=sha256:a2b87058b1a6d47e08a56b97d652f1d727289e226e394bfe84615c1f5cba6cd3

Observation 82e4b745-76e2-4cf8-b463-6c323e630f29 · outbound

This paper cites Parameter-efficient transfer learning for nlp.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Parameter-efficient transfer learning for nlp

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.001674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.549128Z digest=sha256:67fcedc6cd79f6f1888ee8a7f174ae34d47ffe2f7a25c07f9652c5393183b3da

Observation 6625c00c-b459-4507-a833-94e918b277d1 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval LoRA: Low-Rank Adaptation of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.551012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.551012Z digest=sha256:796c83a9ebcc4d6f82af7280a4c0208a78be42c49d83311f9d46d225b3673b0a

Observation 0329a248-4f29-4b20-8274-a3a2a8f55253 · outbound

This paper cites V op: Text-video co- operative prompt tuning for cross-modal retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval V op: Text-video co- operative prompt tuning for cross-modal retrieval

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.994442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.553091Z digest=sha256:d4e2fdbb4987f4b177023ce8b188745dd81b91bf6cbc9d5b327d35f716777ecf

Observation 1c1c649c-d85c-4583-90dc-8678cbbb7fbd · outbound

This paper cites Vi- sual prompt tuning.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Vi- sual prompt tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.987864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.555136Z digest=sha256:5b626b1b9e474bef91de17f9266dba23b96af6278bbf5de2e519f930e610ff1e

Observation e48bacc2-7389-4021-ac23-3c6a801563a9 · outbound

This paper cites Video- text as game players: Hierarchical banzhaf interaction for cross-modal representation learning.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Video- text as game players: Hierarchical banzhaf interaction for cross-modal representation learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.981161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.557171Z digest=sha256:4cfcbb0c76d8d545c4de36f7a126318a59b517608f9f2a31fdc13f90cdde349f

Observation c79fb09d-18d4-46c3-8db5-a85b3247de15 · outbound

This paper cites Mv-adapter: Multimodal video transfer learning for video text retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Mv-adapter: Multimodal video transfer learning for video text retrieval

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.973336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.559599Z digest=sha256:cb6407464334d7e745ab4a76183a13e12542836bcaec4dbbfee78a85c4cda34f

Observation 903f48d1-5eee-4799-b8b3-5b30427eb340 · outbound

This paper cites Deep visual-semantic align- ments for generating image descriptions.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Deep visual-semantic align- ments for generating image descriptions

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.966127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.561792Z digest=sha256:12e124a13cd9a1f03a917322df21273c4c6b455fed5f961e91df13374162b426

Observation 66598c33-e1f8-43e7-8b1e-b96415f926ba · outbound

This paper cites Maple: Multi-modal prompt learning.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Maple: Multi-modal prompt learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.959199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.563869Z digest=sha256:de5938922e559bbc80c5ae6f744672036def200d7acced399402d68a8defda82

Observation e1b60c54-ab64-4a65-b8c8-ddbdc45a349a · outbound

This paper cites Self-regulating prompts: Foundational model adaptation without forgetting.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Self-regulating prompts: Foundational model adaptation without forgetting

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.952290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.566045Z digest=sha256:36c7dcaef3f9673a72a149ffec3358dabc6f02b0059d402db2aeb003825ba61e

Observation be6e81ff-b0a4-4c84-96dd-753df8724479 · outbound

This paper cites Dense-captioning events in videos.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Dense-captioning events in videos

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.945206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.568125Z digest=sha256:c58e3b485a01ca5f91e7a1d24e458056ae47029d4e45d010945464c74fc01e88

Observation 37d14fd7-c852-4144-b417-c8a86a2965dd · outbound

This paper cites Courier Corporation, 1997.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Courier Corporation, 1997

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.570005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.570005Z digest=sha256:2858e88b80a3a0970f0a6813c6f0e5ab593fdb545b6b12e28e632c8ab957f8a9

Observation 0a4cfc7b-c815-4033-bf1b-934abdf1bb8f · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.932510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.571998Z digest=sha256:bb7cf5926e8c68de53e23ad9292d11ba54576b02f32d2e2d4e7c052ef936ce39

Observation 92789e07-17c7-4199-bb2a-94da68badea5 · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in Neural Infor- mation Processing Systems, 2021.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in Neural Infor- mation Processing Systems, 2021

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.925580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.573993Z digest=sha256:9910b63788b75612e12e0c9b79578bc396179fb126095cfa11995de90493c5f0

Observation 6669b5e9-2724-497d-9548-25bcd1d9e577 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.576455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.576455Z digest=sha256:a9dfae6e98816bff45e8292abf0cd98ce5bb335a9b5048781a473290c1dba3f2

Observation 472df5ca-6f3a-4ae8-8f94-535037418fd3 · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Unmasked teacher: Towards training-efficient video foundation models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.914431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.578756Z digest=sha256:fa12957201637dff22c814c28b3b463e2be8aad500b3452b0cbad8e57d172f17

Observation 2470c477-442b-4551-8d82-6811e1f69e24 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.907185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.580758Z digest=sha256:6499532bc182a655e256dbc45a3effdbe07a7b1a47b22dd803903459c7f5d70e

Observation c4330bd7-a4c8-4b56-b714-198727c40189 · outbound

This paper cites Sgdr: Stochastic gradient descent with warm restarts.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Sgdr: Stochastic gradient descent with warm restarts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.900181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.583189Z digest=sha256:6bb3364d4b7fbca3ee950c5a52f0f3b79ac4cc1bbd21594e68cc22786df2f313

Observation e6db3627-e858-42d0-b83d-3907c87c4145 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neu- rocomputing, 2022.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neu- rocomputing, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.893052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.585284Z digest=sha256:f7bcb4ff2786e992d918807ee96bbe44d3b917b2e18e17473ada59e2afb2eaa5

Observation d1a8c765-90e9-45e0-8f5a-0a79c9e6f4c4 · outbound

This paper cites Ea-vtr: Event-aware video-text retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Ea-vtr: Event-aware video-text retrieval

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.886181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.587649Z digest=sha256:805de32ae782a3251c86f5b4d8307f9f4bc041f1a13a1fb02367690c39af48c3

Observation 675359d4-1a39-4b9f-b23d-8ee148482d73 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.879247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.589815Z digest=sha256:f2e07984248a2a42674ded931b3815ab3f9eb864c25d064f2c8e286a24951f6f

Observation cb69a46c-7eb6-4721-ab43-ac139614fd12 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Representation Learning with Contrastive Predictive Coding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.591729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.591729Z digest=sha256:7bc7447558d1c00838f21d6eafe0c56a4e9c3d236c2dec10538ee046d23c59ef

Observation e4603b17-f9ba-4280-bf00-2e3f29eb956a · outbound

This paper cites Language models are unsu- pervised multitask learners.OpenAI blog, 2019.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Language models are unsu- pervised multitask learners.OpenAI blog, 2019

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.594112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.594112Z digest=sha256:8011ce6f0e9c53b8b8a9b1cb14dc8277516473e584515aae71c1b417ef511683

Observation f5a8f1e4-476f-4dbb-aadf-fa421504ecd9 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Learn- ing transferable visual models from natural language super- vision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.868194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.595898Z digest=sha256:d3f204447b5b50e95424181d250bcfe517accb5035636cc71250f1ad246d583d

Observation 3ddc66e8-4b1c-40f1-a540-51cc0dec58eb · outbound

This paper cites The long-short story of movie description.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval The long-short story of movie description

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.859860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.597925Z digest=sha256:e564007b00a7ae7d028746e063d4be8b479171eaee332e5017134396af948ffd

Observation 20943e42-6839-4b62-ad48-600cbe32f003 · outbound

This paper cites TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.599960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.599960Z digest=sha256:42f559a46644d05608c3b91dab1aa7f1f3c4001ef17baa28329ba772827883f7

Observation 5050387a-a682-48d0-87e4-b2319976e873 · outbound

This paper cites X-reid: Cross-instance transformer for identity-level person re- identification.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval X-reid: Cross-instance transformer for identity-level person re- identification

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.852437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.602395Z digest=sha256:6108696974b11defa7aa115b2ef5bbc36106bae7925d81ebddb6b3d1e08fc7e3

Observation a487edaf-4706-4cd6-8dc1-127138457d33 · outbound

This paper cites Zerocap: Zero-shot image-to-text generation for visual- semantic arithmetic.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Zerocap: Zero-shot image-to-text generation for visual- semantic arithmetic

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.845005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.604375Z digest=sha256:39c8be0cb28d7ccd21e99844e1897c8eb75048bdb6593d57855e1ce342813f63

Observation f644fda7-d10d-4b96-8176-32ee03ac7e91 · outbound

This paper cites Yolov10: Real-time end-to-end object de- tection.Advances in Neural Information Processing Systems,.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Yolov10: Real-time end-to-end object de- tection.Advances in Neural Information Processing Systems,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.837245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.606186Z digest=sha256:85623b09f5384419d78b2705edcd8495f12223b195f6c3a3fa3ef2f172544180

Observation aaef0ede-80b5-4deb-9093-06115b7bdf9d · outbound

This paper cites Text is mass: Modeling as stochastic embedding for text-video retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Text is mass: Modeling as stochastic embedding for text-video retrieval

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.828037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.608355Z digest=sha256:ce526fefb8d0acfe3efe14f767b75ab01cd08f35f5d8f83be595238ca9070ce7

Observation 318d79f3-c194-4bb6-a2ea-277328ae9909 · outbound

This paper cites Disentangled Representation Learning for Text-Video Retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Disentangled Representation Learning for Text-Video Retrieval

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.610482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.610482Z digest=sha256:97c297957364a8876a5f768e639dc987bd4ed08f57cdf9738c62c0d562bc425d

Observation 300d6d22-2ed4-4a14-89bc-5e3f9fe69607 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.612750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.612750Z digest=sha256:f3e931745ab2cc73aedf3b8e678f8469acba8216f6a33aebd139836a685107c7

Observation 13d9f64f-a526-4fb3-8666-890cc69a6917 · outbound

This paper cites Cap4video: What can auxiliary captions do for text-video retrieval? InProceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition, 2023.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Cap4video: What can auxiliary captions do for text-video retrieval? InProceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition, 2023

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.820295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.614769Z digest=sha256:092de2bfc15c8e7efaa3c69bba75b379913a3c447a313312fa59bcc5fb24cc02

Observation d3bb257c-fec5-4d6a-b455-d1824b8df63d · outbound

This paper cites Demystifying clip data.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Demystifying clip data

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.810675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.616890Z digest=sha256:bb206f630bc2b84df992e3900f73d7cd9cb21d1cd571a70cbae7bcb43835b102

Observation 62a76fcd-6b96-400d-a602-ea4d4ecd08fa · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Msr-vtt: A large video description dataset for bridging video and language

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.802277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.618838Z digest=sha256:4798b0426fe720b8cc6df96114dc79b443a211f45f3c3e56ccc5ed0f0adede44

Observation a2d4fc0f-88d0-4624-877e-0ea3b83a0e31 · outbound

This paper cites Show, attend and tell: Neural image caption gen- eration with visual attention.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Show, attend and tell: Neural image caption gen- eration with visual attention

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.794086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.620852Z digest=sha256:cd764432e4a037889cae9ff16597fa4006865c4dcaa3d9680a5729a6055eed20

Observation ac8193da-961f-4068-b1fb-a10e92076d7c · outbound

This paper cites Clip-vip: Adapting pre-trained image-text model to video-language alignment.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Clip-vip: Adapting pre-trained image-text model to video-language alignment

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.785496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.623017Z digest=sha256:0c3aca530412fca0cb9378dc67141d1448fc23e652807987bc6f8c47cea61fc4

Observation 4429bc67-e6b9-413d-8cfc-0eeb2161aa64 · outbound

This paper cites LLMI3D: MLLM-based 3D Perception from a Single 2D Image.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval LLMI3D: MLLM-based 3D Perception from a Single 2D Image

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:04:38.682071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.625078Z digest=sha256:6025442537ef426959da82dde842f019d452c658ffad0f3fb5e7bfa20e55bff7

Observation c7252e5e-663e-4109-b549-6fbcafb3c3dd · outbound

This paper cites HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.627382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.627382Z digest=sha256:a491dd0905c1560a811641ee28399f965749fa62d03a9dc5383150f00e94e71f

Observation cbdea981-e643-4f0b-9a62-f98e4d62c816 · outbound

This paper cites Dgl: Dynamic global-local prompt tuning for text-video re- trieval.Proceedings of the AAAI Conference on Artificial Intelligence, 2024.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Dgl: Dynamic global-local prompt tuning for text-video re- trieval.Proceedings of the AAAI Conference on Artificial Intelligence, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.777722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.629860Z digest=sha256:7c57c53654ed90b5cf4e734c27d12eb08821b9bf2aede03a90225f33f1b439cf

Observation e338126f-3b20-4a2c-813f-89a4f1771003 · outbound

This paper cites Cross-modal and hierarchical modeling of video and text.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Cross-modal and hierarchical modeling of video and text

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.770886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.631911Z digest=sha256:888bc9d27988f45c65f40ccd415f1cef26efb44b339a4e5e2fe5563a9b508040

Observation d77fac82-44a4-4399-8190-8b237b154d11 · outbound

This paper cites Neural Prompt Search.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Neural Prompt Search

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.634022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.634022Z digest=sha256:595ce52100078f337a3292bf80f5345957aabdf7e9b40a48eda3aadaf9fd2121

Observation 7e438381-88ec-420d-8564-76c6862d6121 · outbound

This paper cites Conditional prompt learning for vision-language mod- els.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Conditional prompt learning for vision-language mod- els

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.763763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.636387Z digest=sha256:0b72535dc2179184c9ec85ec40e9909e484ad8536701b77f6c07f8d29eae83f8

Observation 60c6db4a-d10a-4ccb-b7c0-3814677297bb · outbound

This paper cites Learning to prompt for vision-language models.Inter- national Journal of Computer Vision, 2022.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Learning to prompt for vision-language models.Inter- national Journal of Computer Vision, 2022

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.755584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.638481Z digest=sha256:8b80157d6fbdb911c1e571c859d2df53168bdbfdcddb25b32723dbc08a78f46b

Observation 207f0633-d482-4615-9278-a64b4d1f201a · outbound

This paper cites 14,αandβare set to0.3and1.0, respectively.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval 14,αandβare set to0.3and1.0, respectively

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.747510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:04:38.641568Z digest=sha256:2090a3cb6b6b10deeae39ccd29e06c763e55bb45481457bdfb58fed1720b65fb

Pith citing papers

No inbound Pith citation observations are available.