Pith. sign in

Paper Citation Record · LEDGER

Expertized Caption Auto-Enhancement for Video-Text Retrieval

As of 23 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 2 inbound Pith citation observations for arXiv:2502.02885.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02885 v3

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:51:29.261016Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:49:03.038872Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:41:06.814014Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43ea0d3f-6a88-400a-8816-1faee030fb74 · outbound

This paper cites Un- masked teacher: Towards training-efficient video foundation models,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Un- masked teacher: Towards training-efficient video foundation models,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.696729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.122961Z digest=sha256:010682f4c0ff1bb9164269f90fb35f596aa270abe0f264445b6b7d4b3ea21256

Observation e24eff71-66be-46ee-a89f-960ce4464e67 · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.685915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.127239Z digest=sha256:0ac09d001252f3ba937ff7988b0f2601542378bfe5fd0205521d136367090e82

Observation fbc34ced-5b6b-44ba-9302-d687cb966a61 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.674886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.131271Z digest=sha256:a837c7aa138b0991909617726e93d5e4159181c41c16a50f95ab5f5679ccac14

Observation d885a4a9-5386-4a2a-8d18-2715cd315e63 · outbound

This paper cites Clip-vip: Adapting pre-trained image-text model to video-language representation alignment,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Clip-vip: Adapting pre-trained image-text model to video-language representation alignment,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.664590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.135065Z digest=sha256:5b09c1c15ae39587c73d491f6820e9b27a4eb3e992db71e95c5fb779d14f9194

Observation 084cc9dd-47b9-45a1-8ea7-a539e7277e10 · outbound

This paper cites Long-form video- language pre-training with multimodal temporal contrastive learning,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Long-form video- language pre-training with multimodal temporal contrastive learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.654191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.139048Z digest=sha256:6113eab396c5a0858fce7d85d95ee31b0e395e2dde8eec1c368924d6a8d795db

Observation a78c93a6-f973-4b2f-9e37-08aa319a33ae · outbound

This paper cites Litevl: Efficient video-language learning with enhanced spatial-temporal mod- eling,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Litevl: Efficient video-language learning with enhanced spatial-temporal mod- eling,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.643321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.143114Z digest=sha256:70fbf474137b727b462d3070b9eea3794990dd22d5f73300c0ed2558790449c7

Observation 4b77a9ed-dfc2-457b-b4ea-295080229ff5 · outbound

This paper cites CLIP2TV: Align, Match and Distill for Video-Text Retrieval.

Expertized Caption Auto-Enhancement for Video-Text Retrieval CLIP2TV: Align, Match and Distill for Video-Text Retrieval

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T10:51:29.147075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:51:29.147075Z digest=sha256:8c469fde38b440a1d22d386b3f56fec031458d99ba2df29c957e3f94c14ad4e1

Observation a78511e2-5632-4827-841d-41e5e5b8285a · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Less is more: Clipbert for video-and-language learning via sparse sampling,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.633102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.151198Z digest=sha256:07af19a2273362742d1b1c142ed7c789c4e917faaef5650669cc5e3e7566ce8d

Observation 20f71575-29c3-45a2-bc5c-663fd5312806 · outbound

This paper cites X-clip: End-to-end multi-grained con- trastive learning for video-text retrieval,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval X-clip: End-to-end multi-grained con- trastive learning for video-text retrieval,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.622543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.154898Z digest=sha256:7e1cecbc6dd4b0d7e17f18522f30bf12c81d278485c3038103fbaef0609912fa

Observation 62e86228-f5fd-47bd-80ff-50fe959a3636 · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Advancing high-resolution video-language representation with large-scale video transcriptions,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.612035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.158633Z digest=sha256:ef45d2e06c673242b8a3720853b5e5254c1a78a8ade2c56f7e18521f68b8ed2a

Observation bb633cea-6a1b-40ce-a1a7-d7509da2db60 · outbound

This paper cites Bidirectional cross-modal knowledge exploration for video recognition with pre-trained vision-language models,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Bidirectional cross-modal knowledge exploration for video recognition with pre-trained vision-language models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.601539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.162492Z digest=sha256:a754b3f636b9570fd19d755aeebf837c0d0329dd708b0daf38bd0cfeb41f7db9

Observation 5064f4a4-2741-41f2-89d4-e4df90d2c81a · outbound

This paper cites She-net: Syntax-hierarchy-enhanced text-video retrieval,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval She-net: Syntax-hierarchy-enhanced text-video retrieval,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.591079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.166800Z digest=sha256:b3105d8c69998ff076a7512956d673ce0d0cf276d702786ccf42a326934a857a

Observation ceec9b38-2584-4c6e-b585-bf67748092d1 · outbound

This paper cites Teachtext: Crossmodal generalized distillation for text-video retrieval,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Teachtext: Crossmodal generalized distillation for text-video retrieval,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.580904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.170475Z digest=sha256:7b44250c6c6a2b139f0bc57d4e79cb50219c4fb1219dfc80c20d3ae1b505c8db

Observation 4c3a5699-381f-46f9-95e4-016a677f26a3 · outbound

This paper cites Text is mass: Modeling as stochastic embedding for text-video retrieval,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Text is mass: Modeling as stochastic embedding for text-video retrieval,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.570628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.174312Z digest=sha256:8959265acfc3497a7c75b40baf19bef94124145b4fb020699f8b285b2450736c

Observation 3d7c1bb0-8baa-4bf0-9706-2bb8e2502ec9 · outbound

This paper cites Cap4video++: Enhancing video understanding with auxiliary captions,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Cap4video++: Enhancing video understanding with auxiliary captions,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.559369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.178387Z digest=sha256:9c6b777cffe255fca0abab1d31d5be6e930c99f236bd1223eccf4c8319e77765

Observation 8db09856-898d-47d9-8897-847583b1d86b · outbound

This paper cites The neglected tails in vision-language models,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval The neglected tails in vision-language models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.548175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.181778Z digest=sha256:48976feef3a5c106e885b8bb6668b20f10e9c7730a1e8e0a8c5997a7128e412a

Observation 9a71ca16-5362-4b0d-9687-fc31138a502a · outbound

This paper cites Verbs in action: Improving verb understanding in video-language models,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Verbs in action: Improving verb understanding in video-language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.537140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.185341Z digest=sha256:a45ca5b9a602ba76d52f2bf69a08eb7c727954caf2c7c42871c06ead22954589

Observation 67a74f46-f828-4c84-8e47-d1ccd88e212b · outbound

This paper cites Videocon: Robust video-language alignment via contrast captions,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Videocon: Robust video-language alignment via contrast captions,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.525903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.188542Z digest=sha256:f7af8178d0b63e5d8fa6198a3fcd42993d2d7760794ee89b7a4a4cb1c1bdfd97

Observation 2c8ee2f4-1d81-4f21-82a9-86326b8c9edf · outbound

This paper cites DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models.

Expertized Caption Auto-Enhancement for Video-Text Retrieval DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T10:51:29.191631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:51:29.191631Z digest=sha256:6f7c6716564125ef2d6e93bb6c6a771577eab2bd41205a64ba09317bc7bc677d

Observation 7afff887-8912-42b5-8552-76931bf2b82b · outbound

This paper cites Automatic prompt optimization with.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Automatic prompt optimization with

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.514177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.194943Z digest=sha256:99cc07349a7f3842bc3c2b4149bd181bf7dd0f551c3e5e7c58b5366a8d7133fc

Observation 0492c361-dc26-425d-bea4-609ac165b528 · outbound

This paper cites Dynamic prompt learning: Addressing cross-attention leakage for text-based image edit- ing,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Dynamic prompt learning: Addressing cross-attention leakage for text-based image edit- ing,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.502985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.198349Z digest=sha256:2261f98b0c308047759d4a80da706c34f865481d357650d99db49539cf12526d

Observation 18a5d58a-d8dc-486a-bcbf-96dcf678f6e4 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Expertized Caption Auto-Enhancement for Video-Text Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T10:51:29.201639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:51:29.201639Z digest=sha256:941611deba6350891bee615c7b6fe018cb2c4b1154df629b537b4776cc23be19

Observation 850cc9ea-5a95-4a69-9af1-084fa2ceb6fb · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Learning transferable visual models from natural language supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.492215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.205283Z digest=sha256:c150b270a11c0eead45bd632de751ed7bfcc3ceade2f662f3d8a206bbafab7a3

Observation 39fef332-2f04-4f03-87f3-699628cf4cdc · outbound

This paper cites Ts2-net: Token shift and selection trans- former for text-video retrieval,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Ts2-net: Token shift and selection trans- former for text-video retrieval,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.482158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.208352Z digest=sha256:5b69e78f2e31cd6573dfde32c89de3b3931a3470145376a6374655b416595d78

Observation a6b10a50-a0a4-4918-a58f-63ffde2acf0d · outbound

This paper cites Mv-adapter: Multimodal video transfer learning for video text retrieval,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Mv-adapter: Multimodal video transfer learning for video text retrieval,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.471327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.211453Z digest=sha256:2dfff8291eb3e22e21f63705ac38ac35f73ec46f5657d88afc65f7f2c30b3ffd

Observation d60ca170-ae0b-446a-8142-de76b40260b7 · outbound

This paper cites Holistic features are almost sufficient for text-to-video retrieval,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Holistic features are almost sufficient for text-to-video retrieval,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.460553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.214828Z digest=sha256:848f5d0c7259711b22b3d861d9b1b811dc0f57c4ba24982cdcf26b33c150ad28

Observation bb58eef9-2878-47e0-af92-cd278f40d59f · outbound

This paper cites Disentangled Representation Learning for Text-Video Retrieval.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Disentangled Representation Learning for Text-Video Retrieval

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T10:51:29.218350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:51:29.218350Z digest=sha256:f1d4afcd9e07ab144c9bb8475fcf93b703b828905118fd8ff1da8b818be82dc3

Observation c1b0a681-8aec-459c-9e21-b7a8eb9cd8a4 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Expertized Caption Auto-Enhancement for Video-Text Retrieval LLaMA: Open and Efficient Foundation Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T10:51:29.221981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:51:29.221981Z digest=sha256:7265bbec2a4f73f1d7b6d6bf613823761ffc566a0061196476490b1d4ef5f251

Observation c37c38c7-ec0d-4d56-b704-a6a85e782de2 · outbound

This paper cites Reversed in time: A novel temporal- emphasized benchmark for cross-modal video-text retrieval,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Reversed in time: A novel temporal- emphasized benchmark for cross-modal video-text retrieval,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.449360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.225770Z digest=sha256:ec6c3ba6f409722184afe6d2059a89d1c80de6583b214dd997c1258fec7b5cb4

Observation 4789df69-4273-4670-8185-82ec70e6dc28 · outbound

This paper cites Visual instruction tuning,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Visual instruction tuning,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T10:51:29.228981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:51:29.228981Z digest=sha256:9477c959473b364c3d2b0998cd134246ec04319e864a7a8b87f3c1ef5d49737a

Observation e9d3b8af-35c7-40a5-93c1-28f4693e274a · outbound

This paper cites Improving clip training with language rewrites,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Improving clip training with language rewrites,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.432178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.232535Z digest=sha256:bc60e7f00c890791a8bd9dc2625b202f1ddd2ec0cd519e0da09a24204b787da8

Observation 99361dea-d931-4b7d-947d-87196362e4d6 · outbound

This paper cites mplug- 2: A modularized multi-modal foundation model across text, image and video,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval mplug- 2: A modularized multi-modal foundation model across text, image and video,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.421805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.235914Z digest=sha256:8519bf523d45fb0182cfad8f4fe2bf6396d21a7adb505cb07f55e37f08afabcf

Observation a98214e6-1284-46e5-9904-eed78540db68 · outbound

This paper cites Dual-modal attention-enhanced text- video retrieval with triplet partial margin contrastive learning,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Dual-modal attention-enhanced text- video retrieval with triplet partial margin contrastive learning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.409994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.239182Z digest=sha256:573d7b9432b032e89b49d282e62163dc95aa55b40c803b02fc269d46b409fccf

Observation 22e7c9ee-d26f-49db-a4e7-d6b331c666b5 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Msr-vtt: A large video description dataset for bridging video and language,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.399288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.242862Z digest=sha256:d4f5703aff7a0445d34666087ef2d393250bfd30e4c95e8af9198f3afa9f65db

Observation c5b5af83-5b05-4bd8-b152-d263ee9880ce · outbound

This paper cites Use What You Have: Video Retrieval Using Representations From Collaborative Experts.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Use What You Have: Video Retrieval Using Representations From Collaborative Experts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T10:51:29.246066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:51:29.246066Z digest=sha256:2e75c413d44df15278746beb1b491e00fa69021a1ae7af5378092a0f439ddc1c

Observation 566b2468-af15-439a-8062-096dc8610593 · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Collecting highly parallel data for paraphrase evaluation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.388645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.250564Z digest=sha256:d24bc5dd41f39b367cd1024a77922d528a2023665a2001d3023644d5dc6905fa

Observation 70739db6-90e6-4f3f-8833-9c40e6eedc32 · outbound

This paper cites Localizing moments in video with natural language,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Localizing moments in video with natural language,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.377289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.254278Z digest=sha256:524891bfd4b24f6a9ac28e2eb1235cf7d861f9092ec800177b76dc14f940e259

Observation b034008b-686c-4a8c-b0cc-f03527e36b19 · outbound

This paper cites M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation.

Expertized Caption Auto-Enhancement for Video-Text Retrieval M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T10:51:29.257437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:51:29.257437Z digest=sha256:e7c545b51fa993a3abfee477295de504a8d0c2425d69c79a62d2153fef41c2b7

Observation 914bb2f5-da5f-4acd-8c7d-14a79bc37ffa · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning,.

Expertized Caption Auto-Enhancement for Video-Text Retrieval Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T10:51:29.365989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T10:51:29.261016Z digest=sha256:3f463bd53b362bbd1c3bb3182c4cf4b1b25c77b04615126b45844ef0a2de6a64

Pith citing papers

Observation c435583e-9907-42d8-9e71-7104318873c8 · inbound

MemVerse: Multimodal Memory for Lifelong Learning Agents cites this paper.

MemVerse: Multimodal Memory for Lifelong Learning Agents Expertized Caption Auto-Enhancement for Video-Text Retrieval

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T18:49:03.038872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:49:03.038872Z digest=sha256:a067265b895d82322066c288178e7da51bced15c0f089b01b82f4026fd617d62

Observation 721cd6cd-30bd-4b95-9f61-3a8e8471d1b0 · inbound

EmbodiedGovBench: A Benchmark for Governance, Recovery, and Upgrade Safety in Embodied Agent Systems cites this paper.

EmbodiedGovBench: A Benchmark for Governance, Recovery, and Upgrade Safety in Embodied Agent Systems Expertized Caption Auto-Enhancement for Video-Text Retrieval

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:06.821330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:21:20.231759Z digest=sha256:da79092685c034af74b09c14a4327f3ee6412c94364a00d49c8b76a6389d1946