Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:27:33.138881Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2510.15470.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:27:33.138881Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b0c0aa15-cc85-4c71-a286-484f31741ef5 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Overview and current status of remote sensing applications based on unmanned aerial vehicles (uavs),
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a324bb2-db7c-440f-824b-071e0ab3a1d3 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Unmanned aerial systems for photogram- metry and remote sensing: A review - sciencedirect,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dc5533e-70c2-4138-9b5c-7f339c05fd3e · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Uavs challenge to assess water stress for sustainable agriculture,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8765432-905a-40af-8f8a-1a490c021249 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Sustainable agriculture by increasing nitrogen fertilizer efficiency using low-resolution camera mounted on unmanned aerial vehicles
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e896907-0e14-4f67-8316-2bcf1ac1e3ba · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90b5870b-ae9f-4d53-ac0c-6c1efecb6d73 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Visible-thermal UA V tracking: A large-scale benchmark and new baseline,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b1acb93-97e2-4dce-84b4-eaff330f7baa · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval High-resolution feature pyramid network for small object detection on drone view,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43aab4db-a5c4-463e-bdf8-2b815cc304d3 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval EarthNets: Empowering AI in Earth Observation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8e3df58-a671-41e1-b783-544495bf167e · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Sdanet: Semantic- embedded density adaptive network for moving vehicle detection in satellite videos,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5ee1637-38b1-4c68-bf0c-e4ad8138e9e0 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Pareto refocusing for drone-view object detection,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4436b924-fa26-430b-96de-42e4be59cc48 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Temporal-spatial feature interac- tion network for multi-drone multi-object tracking,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c74a6b5-40d5-44cb-b2e5-442fca50f8a0 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Cross-drone transformer network for robust single object tracking,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ef94245-84c3-4cb3-a8e7-e4fe1b743173 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Transformer- based spatio-temporal unsupervised traffic anomaly detection in aerial videos,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9d4e355-cd8d-4900-8728-6c95121940a6 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Visual contextual semantic reasoning for cross-modal drone image-text retrieval,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9441b470-f639-4d38-9f0a-322d8002046a · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Deep saliency smoothing hashing for drone image retrieval,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bb93069-d2d8-40de-92e8-c24abffa9d10 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Uav-human: A large benchmark for human behavior understanding with unmanned aerial vehicles,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47de33d5-058a-4de1-ad42-b20c3d5df04c · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Multi-modal transformer for video retrieval,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ee3c2bb-ee3c-4523-b65a-165f8d75eb13 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Towards Video Anomaly Retrieval from Video Anomaly Detection: New Benchmarks and Model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da72a306-d5d3-4c31-b58a-599f17bb7e8a · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Multilevel semantic interaction alignment for video-text cross-modal retrieval,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60aa7b8b-24f2-4e87-a95e-39345f62f1d4 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Videoclip: Contrastive pre- training for zero-shot video-text understanding,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93c71b1d-601f-4056-8acf-fcd1f85765aa · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval T2VLAD: global-local sequence alignment for text-video retrieval,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f26efacf-4cd4-448b-b021-45248dea36ad · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d77365a2-c24e-47af-8c97-2dfa1d583892 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Centerclip: Token clustering for efficient text-video retrieval,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3ed3792-c2ec-42b0-9593-42dd6fb4828e · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval X-pool: Cross-modal language-video attention for text-video retrieval,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c5e0c6e-f8a6-4630-a73f-a9b81851972e · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval X-CLIP: end-to- end multi-grained contrastive learning for video-text retrieval,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5c3f83a-82f4-4d47-98b6-37f333a9780c · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Less is more: Clipbert for video-and-language learning via sparse sampling,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9803a14-b0eb-430d-a80d-598ef9db6fcb · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Frozen in time: A joint video and image encoder for end-to-end retrieval,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 951f7b00-d963-43fd-98f4-5b50d3095524 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Align and tell: Boosting text-video retrieval with local alignment and fine-grained supervision,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2c5b1f2-781e-4669-9972-6dbc07a29a33 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Ts2-net: Token shift and selection transformer for text-video retrieval,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f527a65-d759-4557-84a5-aa1fdfbb1fb7 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee628d70-cb9f-4f04-8bbf-bb9e9d5be3f5 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Modeling uncertainty with hedged instance embeddings,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f782ef05-09e4-4362-ad43-99079a6f13f7 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Probabilistic face embeddings,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6bab12d-53b6-4083-9e3a-1ab732218cdd · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Data uncertainty learning in face recognition,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb41864a-658f-4ba8-8766-69d9eb7510e5 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval View- invariant probabilistic embedding for human pose,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 873e0b6f-5d08-406b-bf91-ed57d0c74aef · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Probabilistic embeddings for cross-modal retrieval,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa5eb562-b6de-4f19-8e39-63fa6dc94db5 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval UATVR: uncertainty-adaptive text-video retrieval,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b20c5da-334b-4528-b4f4-4aed9c1dc539 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4161197-3ad7-4cc3-96fb-2cad8ccf6042 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Position-guided Text Prompt for Vision-Language Pre-training
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55782fe3-41b3-400b-84f8-223174cbb07d · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Multi-Modal Representation Learning with Text-Driven Soft Masks
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8fa8345-6409-4780-8603-142a90d5912b · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval GALIP: Generative Adversarial CLIPs for Text-to-Image Synthesis
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b90c1fea-d017-40a4-8b86-c7722c51034b · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval MAGE: MAsked Generative Encoder to Unify Representation Learning and Image Synthesis
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afae6587-692c-4e98-8b8d-1a7523cd624b · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Towards accurate text-based image captioning with content diversity exploration,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bd18602-762d-4dad-9fc4-9408ba13c2ab · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Show and tell: A neural image caption generator,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18f86d71-9ce3-48b6-b2e6-24ba3afd280c · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Deep visual-semantic alignments for generating image descriptions,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4c571ea-4557-4e20-be37-90c37aa54e39 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Visual semantic reasoning for image-text matching,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 872d62fc-edf7-4732-b12e-326a1ae016c6 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Conditional prompt learning for vision-language models,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e3137e0-7a29-4ed9-b12a-65283ea13786 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Bridging video-text retrieval with multiple choice questions,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba364f4b-3f3f-4c41-ae53-b72a7cf8fd04 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Visual abductive reasoning,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a46d115a-6ce9-4df6-97c3-439448401aeb · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Membridge: Video-language pre-training with memory- augmented inter-modality bridge,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 923334b6-6923-4ac6-834c-51880dc0bc68 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval A straightforward framework for video retrieval using CLIP,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ff43d9-5430-4dac-8342-1a89a74d838c · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Hisa: Hierarchically semantic associating for video temporal grounding,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65e664f4-4038-4b80-b2f6-896b07ce172f · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Concept-aware video captioning: Describing videos with effective prior information,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 002c68af-697b-402a-9c69-ac0050ef0f40 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Hierarchical representation network with auxiliary tasks for video captioning and video question answering,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d73c0d86-0663-4861-82f5-e8790bd60c9a · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Cross-attentional spatio-temporal semantic graph networks for video question answering,
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 420785a2-87ff-4ebc-9995-d7fa2e27604e · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Adaptive spatio- temporal graph enhanced vision-language representation for video QA,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a1d09eb-1f79-4419-a6dc-d1e6056ac403 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Exploring language hierarchy for video grounding,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d75c092c-1544-43d0-a39e-b17988689ab7 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Layer Normalization
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66a2ca70-a2d6-4263-b222-c076b00d3a6c · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Attention is all you need,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a18ce8eb-18a6-408f-8d67-a1f5f22e59de · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval ERA: A Dataset and Deep Learning Benchmark for Event Recognition in Aerial Videos
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68194454-ded6-4e32-918d-2dfbe69aee6b · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 470f4a6b-2dfe-4047-8fdb-ddbf8073072c · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Decoupled weight decay regularization,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f834c26c-f347-4705-9ba6-bb0f73c85096 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval SGDR: stochastic gradient descent with warm restarts,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75667931-de7d-47aa-a09f-601be6a4bd2d · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Disentangled Representation Learning for Text-Video Retrieval
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aef1d775-333c-4235-b609-b219180a5b23 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Unified coarse-to-fine alignment for video-text retrieval,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7db79281-f41a-4fcb-9d0b-321f8338fdf3 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Text is MASS: modeling as stochastic embedding for text- video retrieval,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e4f03fb-4a4a-41b7-b328-e08aab1cbf7c · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval DGL: dynamic global-local prompt tuning for text-video retrieval,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0777012d-2953-4d31-9fae-33978252ff21 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Text-video retrieval with global-localsemantic consistent learning,
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1a50d70-ce7a-42e3-ae7c-798b52a53934 · outbound
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Tempme: Video temporal token merging for efficient text-video re- trieval,
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.