Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:19:58.847247Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.08590.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:19:58.847247Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e4cc6cc2-5ebc-46f6-9f6f-bdcaf9a95f63 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Faster r-cnn: Towards real-time object detection with region proposal networks,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 820545a0-8f94-4934-b661-03f81fcd31c3 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Stacked cross attention for image-text matching,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cdfd30ec-1d79-4a7c-be7b-f4ca4a409cfe · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Image- text embedding learning via visual and textual semantic reasoning,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 514a9348-726f-42c5-b9a6-7cd810702ce2 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching At- tentive mask clip,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c5fe10b2-5c13-4113-89f0-2502927467bb · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Composing object relations and attributes for image-text matching,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8d94c9a9-67e6-475f-adf3-a00f85950f4e · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Learning transferable visual models from natural language supervi- sion,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8ca9b18e-f722-4b92-b8a7-a6d85806f3d2 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Sigmoid loss for language image pre-training,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2f68182e-5c82-4a45-bea4-6ba588a05ed2 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Regionclip: Region-based language-image pretraining,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 30654223-7950-4359-9836-f79e4b7a9526 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Sclip: Rethinking self-attention for dense vision-language inference,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9bccf390-e917-45ae-b837-fc8e5ac865aa · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Improving clip training with language rewrites,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f1acdda5-029f-4367-ad21-f8a58209a01a · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Improving multimodal datasets with image captioning,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cbb654d8-1fa7-4e50-b034-6b2ba87c5c8a · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0ee958d2-bfaf-42c2-be02-df1ebbc15013 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching MLLMs-Augmented Visual-Language Representation Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6af04eb6-4472-4714-bb0f-7cc3062c3059 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0033427b-733e-4acb-b098-13d0f736b924 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching C-Pack: Packed Resources For General Chinese Embeddings
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08f96dca-8f4c-404a-be87-e9c1a9a24f75 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Fine-grained image- text matching by cross-modal hard aligning network,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 50e81f65-4615-40ad-b95f-ed89de7a1fc4 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Fast, accurate, and lightweight memory-enhanced embedding learning framework for image-text retrieval,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fcd0442a-4ab7-452f-a98d-8230c52abf65 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Learning the best pooling strategy for visual semantic embedding,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e001793f-3add-4bdf-9b6a-1e4239ba249c · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Learning semantic relationship among instances for image-text matching,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 26a67493-2443-427f-b1e0-30aa4136f551 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Sinkhorn distances: Lightspeed computation of optimal transport,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f2d0d17a-2ff4-4404-9b24-cbaf2c7c7a03 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Deep visual-semantic alignments for generating image descriptions,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 593e71a4-bd87-4236-a740-ca1389172458 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching N24News: A New Dataset for Multimodal News Classification
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e83deb29-8974-4017-8f07-bfd8cee1eefa · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b76fc5ee-b355-4008-8d7e-2ea5b2803cc3 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Florence-2: Advancing a unified representation for a variety of vision tasks,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 337190bb-dc06-4ccb-9f17-3e968b6363a9 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Align before fuse: Vision and language representation learning with momentum distillation,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8ab328d0-2d7d-4a6b-bbe6-4b9ed6e2c7ca · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Remote sensing cross-modal text-image retrieval based on global and local information,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5d951393-c6d2-46c2-b4a3-189abee9d957 · outbound
Visual Semantic Description Generation with MLLMs for Image-Text Matching Hypersphere-based remote sensing cross-modal text-image retrieval via curriculum learning,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.