Pith. sign in

Paper Citation Record · LEDGER

Visual Semantic Description Generation with MLLMs for Image-Text Matching

As of 10 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.08590.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08590 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:19:58.847247Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4cc6cc2-5ebc-46f6-9f6f-bdcaf9a95f63 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Faster r-cnn: Towards real-time object detection with region proposal networks,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.439282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.647447Z digest=sha256:6181fff065a7550e96c5057ca87e233b56e8a3afad4a9c8d47fe706869ae95f3

Observation 820545a0-8f94-4934-b661-03f81fcd31c3 · outbound

This paper cites Stacked cross attention for image-text matching,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Stacked cross attention for image-text matching,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.421214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.652953Z digest=sha256:abcbf040b138adab6106b6f60e9e3815607fd4ae60842e205871831b576163aa

Observation cdfd30ec-1d79-4a7c-be7b-f4ca4a409cfe · outbound

This paper cites Image- text embedding learning via visual and textual semantic reasoning,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Image- text embedding learning via visual and textual semantic reasoning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.397910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.657786Z digest=sha256:8c9dffc20561dc966a49f220731bd58b31241bd49473edd949799c47dcaf765a

Observation 514a9348-726f-42c5-b9a6-7cd810702ce2 · outbound

This paper cites At- tentive mask clip,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching At- tentive mask clip,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.367248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.663498Z digest=sha256:f956eec7bfd3b016e0736c041de30f7582a4d51ce7e02ce2925d197c2a97bf45

Observation c5fe10b2-5c13-4113-89f0-2502927467bb · outbound

This paper cites Composing object relations and attributes for image-text matching,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Composing object relations and attributes for image-text matching,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.352015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.669323Z digest=sha256:a2b2b0d9e26f2d02402cb4e6b6c56c1cd959c45897c4fad244052102e0b44eb9

Observation 8d94c9a9-67e6-475f-adf3-a00f85950f4e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Learning transferable visual models from natural language supervi- sion,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.336427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.675803Z digest=sha256:d0c8905872a0cfeaabd5d6a510b06079cf9982e12bb3a460e2bce3d914467741

Observation 8ca9b18e-f722-4b92-b8a7-a6d85806f3d2 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Sigmoid loss for language image pre-training,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.319859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.683871Z digest=sha256:6ec6f3289db4c416d2796051dd16261a76e7d249c841e8ce122b7e50c44a7234

Observation 2f68182e-5c82-4a45-bea4-6ba588a05ed2 · outbound

This paper cites Regionclip: Region-based language-image pretraining,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Regionclip: Region-based language-image pretraining,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.292337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.693360Z digest=sha256:253a1d2f9178ee2970f8b81154c4e0cba68a6a67759e0f0ad1745c41382b0186

Observation 30654223-7950-4359-9836-f79e4b7a9526 · outbound

This paper cites Sclip: Rethinking self-attention for dense vision-language inference,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Sclip: Rethinking self-attention for dense vision-language inference,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.267034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.703590Z digest=sha256:75288809c25489744ae5a591da3f48177cba942745712c2641eb4fefc327eea0

Observation 9bccf390-e917-45ae-b837-fc8e5ac865aa · outbound

This paper cites Improving clip training with language rewrites,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Improving clip training with language rewrites,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.250030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.715883Z digest=sha256:114347ca4152578e47cd7eeb6c353b0e80469d27ac1212ad31dc7f6b692282e6

Observation f1acdda5-029f-4367-ad21-f8a58209a01a · outbound

This paper cites Improving multimodal datasets with image captioning,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Improving multimodal datasets with image captioning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.234196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.728908Z digest=sha256:b9921ec809bd938c27f0da098db39d72b0a9af2f87a0c5e9e944ef79bf782546

Observation cbb654d8-1fa7-4e50-b034-6b2ba87c5c8a · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.218150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.735122Z digest=sha256:9069e2bff959becf24f58da4ad74c44bc5f8c9c6f3400c141473362ba49f0b88

Observation 0ee958d2-bfaf-42c2-be02-df1ebbc15013 · outbound

This paper cites MLLMs-Augmented Visual-Language Representation Learning.

Visual Semantic Description Generation with MLLMs for Image-Text Matching MLLMs-Augmented Visual-Language Representation Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.740424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.740424Z digest=sha256:632a95fb057f1091510ab94257ca15102f63cef431834d3c86314ae09311fea0

Observation 6af04eb6-4472-4714-bb0f-7cc3062c3059 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Visual Semantic Description Generation with MLLMs for Image-Text Matching MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.748432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.748432Z digest=sha256:d8de2759e6c0ba51285918f8cb2cd63bb4347b7c0cf7f204a6e343e122b38be9

Observation 0033427b-733e-4acb-b098-13d0f736b924 · outbound

This paper cites C-Pack: Packed Resources For General Chinese Embeddings.

Visual Semantic Description Generation with MLLMs for Image-Text Matching C-Pack: Packed Resources For General Chinese Embeddings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.753670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.753670Z digest=sha256:c78ebd460960f554fe5c9e817196f65d637dcba18ce4876f4195bf02643dfad4

Observation 08f96dca-8f4c-404a-be87-e9c1a9a24f75 · outbound

This paper cites Fine-grained image- text matching by cross-modal hard aligning network,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Fine-grained image- text matching by cross-modal hard aligning network,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.201309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.761302Z digest=sha256:d7a0bf1f7395132509ad112be4b38d35e5c11d58138858d467802b4ae8646e0a

Observation 50e81f65-4615-40ad-b95f-ed89de7a1fc4 · outbound

This paper cites Fast, accurate, and lightweight memory-enhanced embedding learning framework for image-text retrieval,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Fast, accurate, and lightweight memory-enhanced embedding learning framework for image-text retrieval,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.183805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.768032Z digest=sha256:ff272886c6ce777177049c3ffa519dfd6ddc1dc04cb13aaf2400161f31c99907

Observation fcd0442a-4ab7-452f-a98d-8230c52abf65 · outbound

This paper cites Learning the best pooling strategy for visual semantic embedding,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Learning the best pooling strategy for visual semantic embedding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.166182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.772689Z digest=sha256:552a5d94ff6e703746af9b5cbf806d6518c6aaa8436e7f46b50226771c2d6f5c

Observation e001793f-3add-4bdf-9b6a-1e4239ba249c · outbound

This paper cites Learning semantic relationship among instances for image-text matching,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Learning semantic relationship among instances for image-text matching,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.145425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.777092Z digest=sha256:8dcf5cacda14f9a751ca15598d31cbebc88279de44454903753b38b7a3ae4059

Observation 26a67493-2443-427f-b1e0-30aa4136f551 · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Sinkhorn distances: Lightspeed computation of optimal transport,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.124932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.797389Z digest=sha256:fe1f941a131d84177feb07eeea2ec1a26fcdaca4ce65c1a9614eed0334ad2720

Observation f2d0d17a-2ff4-4404-9b24-cbaf2c7c7a03 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Deep visual-semantic alignments for generating image descriptions,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.105418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.805931Z digest=sha256:e47c277c5ed588bee217e3410a6dc63c5eee679c9877cdea3df0c15394dec215

Observation 593e71a4-bd87-4236-a740-ca1389172458 · outbound

This paper cites N24News: A New Dataset for Multimodal News Classification.

Visual Semantic Description Generation with MLLMs for Image-Text Matching N24News: A New Dataset for Multimodal News Classification

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.811615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.811615Z digest=sha256:0ae2c9d92da4af45e4520985894d5b366e7d9379b0112250ccdea7c6a959a8dd

Observation e83deb29-8974-4017-8f07-bfd8cee1eefa · outbound

This paper cites Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.084295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.817342Z digest=sha256:97d85090e5c09771bda55775f5c5a9ce35ad8bfe6882b60640f15de2f9eaa61a

Observation b76fc5ee-b355-4008-8d7e-2ea5b2803cc3 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Florence-2: Advancing a unified representation for a variety of vision tasks,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.065552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.825844Z digest=sha256:34e08474215ccd1e41d58278161203925c2340d26b00fb2c354ceab482184072

Observation 337190bb-dc06-4ccb-9f17-3e968b6363a9 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Align before fuse: Vision and language representation learning with momentum distillation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.045119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.832266Z digest=sha256:b8b175ce68a05c20ed9cfc21edba6970e35e04cfe25e25556458ed1911c8d854

Observation 8ab328d0-2d7d-4a6b-bbe6-4b9ed6e2c7ca · outbound

This paper cites Remote sensing cross-modal text-image retrieval based on global and local information,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Remote sensing cross-modal text-image retrieval based on global and local information,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.019781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.842725Z digest=sha256:69e1abe5ca75537868292abc8adea145474c5391f96cf01aeb6b9501e42f28a5

Observation 5d951393-c6d2-46c2-b4a3-189abee9d957 · outbound

This paper cites Hypersphere-based remote sensing cross-modal text-image retrieval via curriculum learning,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Hypersphere-based remote sensing cross-modal text-image retrieval via curriculum learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.000703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:19:58.847247Z digest=sha256:5c05c27824068fea7cacde774ac20903caabaf0c51b7a116eda61cd67da351ae

Pith citing papers

No inbound Pith citation observations are available.