Pith. sign in

Paper Citation Record · LEDGER

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval

As of 10 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 0 inbound Pith citation observations for arXiv:2601.20597.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.20597 v2

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T10:40:01.902254Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

81 of 81 outbound references displayed

  • verified exact2
  • verified fuzzy69
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1116804f-5659-4cc0-9b45-1d9667496262 · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.709794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:e86f7eb24052a1b561a159fdb72b3a0c879b226e00c7899dea47cd1d7709153f

Observation 4f82f5e7-7c65-4c5d-9e40-dee810632c0d · outbound

This paper cites Ashok, K.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Ashok, K

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.714676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:c19fbd9c39797bc2c2aadb0003d163630aaccd3e0cccab1f836d477001892f99

Observation 6e0b33d6-b5fd-43b8-860c-b96508f0c9e3 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.707392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:f0a587ab766ce5ae64a11aff77fffb75faa09ac1d196a5bb597770974b08edcd

Observation 37f15509-ad8c-4d42-87e0-b2800cb4290c · outbound

This paper cites Non-autoregressive cross- modal coherence modelling.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Non-autoregressive cross- modal coherence modelling

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.716828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:201b9b6565e5108c69a1e68c8967d98fa0ac94df3c22402d919d35c85fd00704

Observation 1982bc2e-ebc8-4b69-9931-0a75c947a475 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Activitynet: A large-scale video benchmark for human activity understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.705324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:3cd26dcd76dce2e3966d22aefbf4e215ca76e85779696c5c191855a85009fa9d

Observation fdc776cf-03f6-418a-a7fd-e847250d8444 · outbound

This paper cites Online fast adaptation and knowledge accumulation (osaka): A new approach to continual learning.Advances in Neural Information Processing Systems, 33:16532–16545.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Online fast adaptation and knowledge accumulation (osaka): A new approach to continual learning.Advances in Neural Information Processing Systems, 33:16532–16545

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.709995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:46b2d50889feab23f07140ea5ae50b3033442760cd8c8c18d8312bb6e85d14f8

Observation dcab8dd2-25f9-4a20-9dff-5d406f0fad9e · outbound

This paper cites Clumo: Cluster- based modality fusion prompt for continual learning in visual question answering.Journal of Artificial Intelligence Research, 83.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Clumo: Cluster- based modality fusion prompt for continual learning in visual question answering.Journal of Artificial Intelligence Research, 83

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.712837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:b9bc99fe1151a6be11dc6cfecb278da33ffc29521ebd4cb1b3333f4a4f95bf35

Observation 7681fd1e-64b4-4899-8472-6a0cfa999907 · outbound

This paper cites Castro, Manuel J.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Castro, Manuel J

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.701475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:f44b5ac7221a9d2fad4eacbfd99aa928a825e398f342ffeec1150ddea69afde9

Observation 3b2f0bbb-149a-4ba9-8c8a-369115acf023 · outbound

This paper cites Fine-grained video-text retrieval with hierarchical graph reasoning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Fine-grained video-text retrieval with hierarchical graph reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.694266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:e970d469b53c808c5961e79c45d32f6d8c042208b01958ada83da82b202afeae

Observation da9974f5-19a1-4cfb-9e61-c1b7a53ee94f · outbound

This paper cites Vision- sensor attention based continual multimodal egocentric ac- tivity recognition.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Vision- sensor attention based continual multimodal egocentric ac- tivity recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.703508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:66b7c6fc21b96f70b2faefbcef8aba559b832d55a6db387cfcea368d02b5be2b

Observation 169e921f-6c01-41e7-abb1-4d21d640f6a5 · outbound

This paper cites Teachtext: Cross-modal generalized distillation for text-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Teachtext: Cross-modal generalized distillation for text-video retrieval

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.705551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:5a06d24ca811ec59d740c76043ddd6d4b80837e82af416d06ea8d166bf309ded

Observation 8c686c0e-6338-41a2-9249-55373b401d8d · outbound

This paper cites Don't Stop Learning: Towards Continual Learning for the CLIP Model.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Don't Stop Learning: Towards Continual Learning for the CLIP Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:40:50.904706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:85288c95fe69f3befefb6518906854075cb310455ad44e58b60625b1e57ae558

Observation 227e82ff-e973-409c-b3cb-f3a67b75ef71 · outbound

This paper cites Podnet: Pooled outputs distillation for small-tasks incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Podnet: Pooled outputs distillation for small-tasks incremental learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.696497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:f99a0028d9ee12cb7a9d1bc6aa2030da1dfffbd11a981739453420ea142344ee

Observation d136fbec-b734-4c12-aa50-905633e73a58 · outbound

This paper cites A feature- space multimodal data augmentation technique for text- video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval A feature- space multimodal data augmentation technique for text- video retrieval

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.598654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:89306d0566b8a47ac622661fec9fe35be8367fad04767d2637aea34fa95d0408

Observation cd9416df-fac6-48d9-91d0-cdce30f56e10 · outbound

This paper cites Uatvr: Uncertainty-adaptive text-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Uatvr: Uncertainty-adaptive text-video retrieval

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.600733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:e0f452760e921bc4f8b43f95d4bc2c7adc6b6ed8cdf7d3a13436ae7825421093

Observation c99db1a2-c20f-4648-a484-33184fd37811 · outbound

This paper cites Transferring image-clip to video-text retrieval via temporal relations.IEEE Transactions on Multimedia, 25:7772–7785.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Transferring image-clip to video-text retrieval via temporal relations.IEEE Transactions on Multimedia, 25:7772–7785

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.618257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:5ad3a6ce31f3fa3a8bff13a2d92af113daabd7cbd6d34ec7a3d4c8465a146b77

Observation c0b5fd60-b1d3-4dee-b8fa-af287fae4eef · outbound

This paper cites Multi-modal transformer for video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Multi-modal transformer for video retrieval

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.620652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:369ec7f23cf40f2e5d7121bb13b3d36327658a67523fccf946c853bba0e12351

Observation 52c14299-a180-40fa-93af-00ca30203a9a · outbound

This paper cites X-pool: Cross-modal language-video attention for text- video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval X-pool: Cross-modal language-video attention for text- video retrieval

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.625377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:39fff1b29b4b95ddbf4c5dfee1f2e29fe7a3c2411e5fe4259335357e42202149

Observation ea1e9ff3-71d0-499b-b885-fedd0577eae9 · outbound

This paper cites Dyson: Dynamic feature space self- organization for online task-free class incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Dyson: Dynamic feature space self- organization for online task-free class incremental learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.630233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:0a4cfba7617f9bb9e1d00037ceba0fa558d4e22e1749a135a6076a945e33e793

Observation 235bf544-8942-41bd-9e04-cc4ce4c8aa7b · outbound

This paper cites Learning a unified classifier incrementally via rebalancing.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Learning a unified classifier incrementally via rebalancing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.639797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:6d6d4c05b877da770f2facd59e789eff1ed74880dd3ccd640bfb3c089ff6d1f9

Observation 23e57aa6-d95f-4629-a46d-21e504b01840 · outbound

This paper cites Curiosity-driven class- incremental learning via adaptive sample selection.IEEE Transactions on Circuits and Systems for Video Technology, 32(12):8660–8673.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Curiosity-driven class- incremental learning via adaptive sample selection.IEEE Transactions on Circuits and Systems for Video Technology, 32(12):8660–8673

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.635083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:4fda322a5f2e904121f754c55043bbdb4f6ee83c5e8ad5c1ff8a4f5cb67f9417

Observation e604ae7e-8b2e-4f53-a92d-f059bf6d3837 · outbound

This paper cites Distilling causal effect of data in class-incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Distilling causal effect of data in class-incremental learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.675445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:78418f6308da218f157f5b4f3d3893d7e258bd44e0b9323291ad25e9c31a16d4

Observation 8480298d-d87b-4d0b-b8ea-1a0e5e4af4bf · outbound

This paper cites Neural collapse inspired federated learning with non-iid data.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Neural collapse inspired federated learning with non-iid data

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.609558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:48d6983d04d985fb3598d6622af26a5df1677757bb10d1667e9e4cf153d0f862

Observation 4dfa2c9d-afc7-46a9-8162-c85451acfb4f · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.582087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:d35c54de87a74cbf20fd385c97f08f421e1fe67648df18da75912b484de0f1de

Observation 8b9fcda3-507a-41ba-92c2-0993e30f8931 · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.573147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:c65deb2027c7ac3733ec46611517f0a8c3aaf0c72a57e5f10720d1c989edf98e

Observation ddb1b485-f960-4d80-b661-c6fd3649bfd3 · outbound

This paper cites Hybrid-tower: Fine-grained pseudo-query interaction and generation for text-to-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Hybrid-tower: Fine-grained pseudo-query interaction and generation for text-to-video retrieval

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.567018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:13ffeca4ce341cc3ad5ef5fc3222697a01cb76cf352ea906ccec813413623baa

Observation ab85f800-dd7a-4da5-a188-4f29c50e368d · outbound

This paper cites Bakker, Nicu Sebe, and Michael S.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Bakker, Nicu Sebe, and Michael S

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.569168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:7d8c3440c02849cb004e6c876813ab679e361095c756303a4989fcf6d718992b

Observation afe372b3-fb8c-47ef-b4ee-d415d6312775 · outbound

This paper cites Dynamic integration of task-specific adapters for class incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Dynamic integration of task-specific adapters for class incremental learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.576420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:5ed6d28382e3eaa591689b23d94602482f989f48bce85a1bd7e2181be3282cb8

Observation a2dd7117-829c-4ad6-ba6f-a74700987012 · outbound

This paper cites Multi-modal inductive framework for text-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Multi-modal inductive framework for text-video retrieval

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.658278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:17e66ec8ff91b4678c81a5a13ed40dde24708a6f83e23a62e9ca1010f186fb56

Observation 3bcd1dc6-93b9-493d-8c62-a0ebf4ce32be · outbound

This paper cites Learning without forgetting.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Learning without forgetting

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.663084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:391afddc4ae7c0403c528e794f0dcda266b0accfd7331ff434bc79cc7032bf11

Observation edebad7e-2da5-4a8d-9e83-d68f5976bc11 · outbound

This paper cites Anchor assisted experience replay for online class- incremental learning.IEEE Transactions on Circuits and Systems for Video Technology, 33(5):2217–2232.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Anchor assisted experience replay for online class- incremental learning.IEEE Transactions on Circuits and Systems for Video Technology, 33(5):2217–2232

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.667584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:3fce01a0b3e143e4c7ab227f02b33d7e218f8dc7c778ee2673e5de4985fd5df8

Observation b3b04c33-f98c-4d7e-b06a-1c2d3b4ab5ff · outbound

This paper cites Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.ACM Computing Surveys, 55(9):1–35.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.ACM Computing Surveys, 55(9):1–35

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.687741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:824fc8106525c4060c4688953b2864341e04a4220215ac343f3e0d4153afbd19

Observation 6f9179a4-001e-4107-967b-5669b4a120f1 · outbound

This paper cites Use what you have: Video retrieval using representations from collaborative experts.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Use what you have: Video retrieval using representations from collaborative experts

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.645249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:04bb1135e5be01e491473121f2cfe54adee320cbb79ab6a7a384b30187ecfee0

Observation eb99e289-a5c8-44d6-bc01-1f01e42901f9 · outbound

This paper cites Adaptive aggregation networks for class-incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Adaptive aggregation networks for class-incremental learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.647251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:f27736bb02a3a29daea4d0385ceb034d8f77cd24f211dac87be2ba7971505656

Observation dd4eb987-5ff5-4cc8-8321-8b6609bd0d7d · outbound

This paper cites Ts2-net: Token shift and selection transformer for text-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Ts2-net: Token shift and selection transformer for text-video retrieval

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.640901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:549b53bf7f4403e6b5d35ae44ea1ab2ebc16e1de857c401e05072d32b9f458b4

Observation a06aaf53-a0db-4a1f-8699-a4c9bf11e4e1 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.643129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:0501600335581c388ab24085dea77d24e07ff2cf632c1da58ac59a4c6eaead90

Observation b9c1695b-e4ae-4135-9907-1c55d730415a · outbound

This paper cites X-clip: End-to-end multi-grained contrastive learning for video-text retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval X-clip: End-to-end multi-grained contrastive learning for video-text retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.649523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:6bda42b689b5954229a38a018544dcb51a4ee5ef8656da8cd1e2559089fd808e

Observation cc713bd8-6722-4803-9dcb-d562db5fcfc9 · outbound

This paper cites Packnet: Adding multiple tasks to a single network by iterative pruning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Packnet: Adding multiple tasks to a single network by iterative pruning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.674326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:942e69b62ddf53a2a610c033cd135732dd88d58dd2c8ba8bebc2e993b8242494

Observation 69ef110c-35c8-4964-b321-bec612227c23 · outbound

This paper cites Prevalence of neural collapse during the terminal phase of deep learning training.Proceedings of the National Academy of Sciences, 117(40):24652–24663.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Prevalence of neural collapse during the terminal phase of deep learning training.Proceedings of the National Academy of Sciences, 117(40):24652–24663

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.661076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:a91ceace1ce289baa8cadaaa95a3e53f3283b2c730fad39a24a39da6d7a7ae8f

Observation ca335b31-e5ac-4caf-8464-f3fcb3b3b882 · outbound

This paper cites Prabhu, P.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Prabhu, P

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.591003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:c85c74cd2d3206662c60a20250941602ff6502f0ea565e4917fa478997512fcf

Observation 396540dc-c96c-4a1c-9ffd-0ff774ac9832 · outbound

This paper cites Learning transferable visual models from natural language supervision.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Learning transferable visual models from natural language supervision

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.605292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:eedcc7d6129829126fc6166521f3626126dc1d53d44a9056606583bb1a39d823

Observation c2124e43-fd82-4aff-86a1-65bc36a5367d · outbound

This paper cites De Melo, Benjamin Van Durme, and Rama Chellappa.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval De Melo, Benjamin Van Durme, and Rama Chellappa

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.620155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:22524afbcea28133f4ffb96ecdbbb74e39a55e4b0c87bea04876b472957d1654

Observation db425f6b-aa98-4063-b043-ce955592cef6 · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.591571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:6f27f73962ed7a168b1a8db4640b3041b945fdc209e0c0482758213e99f5df77

Observation 01acc3f3-6af0-4b52-971d-ed6d2bd58277 · outbound

This paper cites Relation triplet construction for cross-modal text-to-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Relation triplet construction for cross-modal text-to-video retrieval

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.661864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:fa8f0676eb7962d063dbd95fc37f05e4bcc899524b0e30ed0c8aaeee843f90e7

Observation 1e61f864-8c9c-446a-8fd5-297b26d80f5d · outbound

This paper cites Spatial-temporal graphs for cross-modal text2video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Spatial-temporal graphs for cross-modal text2video retrieval

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.666346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:672f7722c921e5f801b8cde2f6a239a9210ccc0f91448a7b6411c76724c77885

Observation 1921058a-4e86-402b-8514-441de59d9813 · outbound

This paper cites Learning endogenous attention for 12 incremental object detection.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Learning endogenous attention for 12 incremental object detection

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.648600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:18abbf810f0c20c3d8891155909496d88c2ea25803e8a3c95ae3489b5b2185f3

Observation a58a7f6e-660d-4dac-b416-4111b8a21b09 · outbound

This paper cites Multimodal continual learning using online dictionary updating.IEEE Transactions on Cognitive and Developmental Systems, 13(1):171–178.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Multimodal continual learning using online dictionary updating.IEEE Transactions on Cognitive and Developmental Systems, 13(1):171–178

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.651198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:96e854683c25f2869387ab00b4b1a631bdb771bc50eb685e2f16d016f112cf1c

Observation 3fdd9fa0-b3d5-4eb6-b708-802f7c1f237c · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.669820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:7552772188741a3faed8614c19df1fae1bcc86e2aa55f6973911371f829451da

Observation 8f1985a2-8ea5-4538-ae4f-cee317aeadb9 · outbound

This paper cites Topology-preserving class-incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Topology-preserving class-incremental learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.670974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:c92edef6de793f806fbbc076f4f54c5151123a3e7af0810b73e5dec6bf1a53ac

Observation b7428cfb-c325-49d1-b8a4-7d36a9a4d9cf · outbound

This paper cites new” while consolidating “known.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval new” while consolidating “known

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.682933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:a5506fcb744a56aab090cedf30117b76a1f20b214492d36c51e1a77c60754734

Observation e983ac38-969c-421f-ace1-8e98a8a855b8 · outbound

This paper cites Holistic features are almost sufficient for text-to- video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Holistic features are almost sufficient for text-to- video retrieval

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.615710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:a6a576cdd2c5b73bf15fd74ef9753aa92b598b6fd7ed6d0cd563a18e036bbd15

Observation c3ecd1c6-91d7-406b-82a4-28246964eb74 · outbound

This paper cites Dualcp: Rehearsal- free domain-incremental learning via dual-level concept prototype.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Dualcp: Rehearsal- free domain-incremental learning via dual-level concept prototype

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.636271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:692bf853c1db8b59e17d72b1974ea4ee2453cb8a08657400ef22082494c2b340

Observation 2ef25a4f-1dcc-43b9-b5cd-c8d759a4a279 · outbound

This paper cites Semantic knowledge guided class-incremental learning.IEEE Transactions on Circuits and Systems for Video Technology, 33(10):5921–5931.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Semantic knowledge guided class-incremental learning.IEEE Transactions on Circuits and Systems for Video Technology, 33(10):5921–5931

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.679224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:1a17a256d1b94902c1da798684c1e32db3e64065a59c4683cca068969e78e25e

Observation 5be729ff-f5be-4411-86c5-2238436c0c81 · outbound

This paper cites Non-exemplar class-incremental learning via adaptive old class reconstruction.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Non-exemplar class-incremental learning via adaptive old class reconstruction

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.676623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:26fdf3c30baa2f59eb43f83dc5dce065631a95e9baee35dc63eaceb74743a74d

Observation b24ac00e-0264-4a54-9f7b-0aca9746794b · outbound

This paper cites T2vlad: Global-local sequence alignment for text-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval T2vlad: Global-local sequence alignment for text-video retrieval

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.663904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:c581cbd949f707b62c789166b758e703272f57e84500aef57c98c9fe67ee0143

Observation 2a1936ad-5e4f-43f7-9ece-dedbc41a08e5 · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.609246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:cdf0d8c2fcac1c09f1e275936179d3866d5f6206fff6e84279b09c9b92e788dc

Observation 37d9e1c2-d717-4481-b2c1-fdb53fc4fe02 · outbound

This paper cites Unified coarse-to-fine alignment for video-text retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unified coarse-to-fine alignment for video-text retrieval

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.624099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:fb783d1295f0b9f33f49801a886b8c96bdae03a5f0ed3b5905983b99c21486b7

Observation 72e1b749-ee1a-482f-8266-452eb054661d · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.683137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:02bc267c41756a7cf514cc85a40f8105e130ecc3ea54d8a88aa01a6c189baa04

Observation f599089f-2aec-42db-b05b-dcf40d1976cc · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.681235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:55094de0946b3acfe923a3f7acca34d121031ede624514a0c72c8fc15ed47e1a

Observation 8575d487-c113-46d4-b50c-953a9fb3053d · outbound

This paper cites Striking a balance between stability and plasticity for class-incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Striking a balance between stability and plasticity for class-incremental learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.622234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:1d49e966023d9ca14bd84826b86b4fa205ae7297b4823974c0a701138599ebe0

Observation bf5cbcf8-e89d-44a7-b88a-955fb23350a6 · outbound

This paper cites Large scale incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Large scale incremental learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.628196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:159c6964fa1388d443d808ba319a301485600ca158fc14adefcc88231c1b18cb

Observation 76bf1e52-9479-4f74-ad57-51d64cfcb98e · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Msr-vtt: A large video description dataset for bridging video and language

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.638408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:a140f942456bd4d96c5d82c85ce424f943ec15c442c8a2438009eb4bd781ad75

Observation 323b25a9-ab67-4122-84d9-172ef450bf81 · outbound

This paper cites Clip-vip: Adapting pre- trained image-text model to video-language alignment.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Clip-vip: Adapting pre- trained image-text model to video-language alignment

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.618052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:9cff8c4cd816d8e6fa9b4dc0510407491ae3366ba1e1a0ca55052a14b5c1c830

Observation 3a3e8bb0-5807-46a3-bcf9-5b05ae7d85a1 · outbound

This paper cites Der: Dynamically expandable representation for class incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Der: Dynamically expandable representation for class incremental learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.632232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:c4de7b9e0794474a84ed3e29a5e02d053508724375fd93c8dc424844c3bd4c9b

Observation 293d60e1-c208-42ac-b50e-fd430be27fa1 · outbound

This paper cites Low-rank prompt interaction for continual vision-language retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Low-rank prompt interaction for continual vision-language retrieval

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.685321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:b43bfcb352ca1e1672901cd36a4538d1f764c7d1582ca490b165181904bbc389

Observation e20f0040-64d1-4218-bc02-81aef12cc4ae · outbound

This paper cites Dynamic support network for few-shot class incremental learning.IEEE Transactions on Pattern Analysis and Machine Intelligence.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Dynamic support network for few-shot class incremental learning.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.673508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:e493f54f2c89a93eb8dcb7641581331be2afaf2857996e9385e84c019e37d4e0

Observation e9a3d12a-91c7-4c8e-ab7f-6b50dee6a5ed · outbound

This paper cites Taco: Token-aware cascade contrastive learning for video-text alignment.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Taco: Token-aware cascade contrastive learning for video-text alignment

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.680932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:2509e0566324230cb50e852f8cc4826f94e5788a9f5105bb9902194c533f096f

Observation e30ce547-9178-4dd1-8db6-b74f4d7e132e · outbound

This paper cites Recent advances of multimodal contin- ual learning: A comprehensive survey.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Recent advances of multimodal contin- ual learning: A comprehensive survey

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:40:50.901410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:72776c4c15a1576422e9b0ab27590224e9c896c1b62f85298531e73a9763163a

Observation 254ee880-1289-427b-92b8-a25f65be6876 · outbound

This paper cites Boosting continual learning of vision-language models via mixture-of-experts adapters.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Boosting continual learning of vision-language models via mixture-of-experts adapters

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.659474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:c1ca1091d029e7358e37f717d102d162a5399ae32b4355faecf505f13185ccc6

Observation 08d621ff-ad31-4363-b78a-363b35155a17 · outbound

This paper cites A joint sequence fusion model for video question answering and retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval A joint sequence fusion model for video question answering and retrieval

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.602712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:897a3a4827c4573f79e80288e1af38a08f8528885966ffea934fded9ade59ef8

Observation 30180881-0e62-4084-8fc1-e23f12b15cf6 · outbound

This paper cites Quantifying and narrowing the unknown: Interactive text-to-video retrieval via uncertainty minimization.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Quantifying and narrowing the unknown: Interactive text-to-video retrieval via uncertainty minimization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.611791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:90d2a1dcce4c6e7c4495a4d639149e3e162b5a5cbd3a229172492bc89b97f0b5

Observation b71c631a-5f11-4577-89f4-5fa686a4dc94 · outbound

This paper cites Mpt: Multi-grained prompt tuning for 13 text-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Mpt: Multi-grained prompt tuning for 13 text-video retrieval

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.655518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:c2b53030293f484cb233e71ca098316a5d3e2071123040d2eacd93f6d7457b46

Observation 01de61a7-4be2-4976-a52d-d3be55779950 · outbound

This paper cites Vqacl: A novel visual question answering continual learning setting.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Vqacl: A novel visual question answering continual learning setting

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.602711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:b85714f9a9c30d79072c20b86b33620c6092f067899702ab7d92abd5e97d26a7

Observation ecf872ac-51ad-464d-8f45-838e029c0fbf · outbound

This paper cites Mgsvf: Multi-grained slow versus fast framework for few-shot class-incremental learning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(3):1576–1588.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Mgsvf: Multi-grained slow versus fast framework for few-shot class-incremental learning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(3):1576–1588

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.690142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:43f438a57fcccb64586e91a15b50fef6380a75f604deaae641d17b38173ba375

Observation c018a695-a164-4cdd-a9e8-4969c9698d9b · outbound

This paper cites Centerclip: Token clustering for efficient text-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Centerclip: Token clustering for efficient text-video retrieval

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.672177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:6b34370cdfdcac1af4a56b8d9373a01f6976ac795a10080ff6debef771038980

Observation 5282bee4-e656-420a-bff7-aebc91fd4343 · outbound

This paper cites Continual text-to-video retrieval with frame fusion and task-aware routing.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Continual text-to-video retrieval with frame fusion and task-aware routing

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.651540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:36f0e92aa410b7bfa1716fa19c3b3443d7ecb0eb8361fde7fa2b5b72ef6593d3

Observation 30e8075a-d419-4b5c-8dd1-9e0ea72e66f0 · outbound

This paper cites Preventing zero-shot transfer degradation in continual learning of vision-language models.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Preventing zero-shot transfer degradation in continual learning of vision-language models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.653601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:5a6a7fb1654e3a1b6d5935b4bb7c46167b83892d0ff269f8fd601cca27a55e0a

Observation 0e8f7948-cea6-4260-ad51-2af72f224ec2 · outbound

This paper cites Understanding imbalanced semantic segmentation through neural collapse.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Understanding imbalanced semantic segmentation through neural collapse

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.665178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:2034b6e6ace0427c1999588f3b2a35f4ef9d0dfd215b1d24a1dac48ced868b66

Observation fc92ad12-59cd-4713-ab6a-5ddf29ecbce1 · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.678753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:d68bfcf48877b796294aefa0e8e8ae87696477be4b97d83cb9e2c3598fe224c2

Observation f406e808-57cf-4780-88ec-ec31b8bf628e · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.626212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:f5b5c845e655f02eea4d9d3c8102a3ba9f44ad7e829a0d0134c37e1dcc6ff184

Observation 68c220ba-d509-46e9-b739-cd48f2e79ba2 · outbound

This paper cites Complementarity- aware space learning for video-text retrieval.IEEE Transactions on Circuits and Systems for Video Technology, 33(8):4362–4374.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Complementarity- aware space learning for video-text retrieval.IEEE Transactions on Circuits and Systems for Video Technology, 33(8):4362–4374

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.614174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:82e4fb5f1ae0968aab986b887d24caa7a02d42956973eaff6822ea0501ad5723

Pith citing papers

No inbound Pith citation observations are available.