Pith. sign in

Paper Citation Record · LEDGER

Visual Semantic Description Generation with MLLMs for Image-Text Matching

As of 8 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.08590.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08590 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:19:58.847247Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4cc6cc2-5ebc-46f6-9f6f-bdcaf9a95f63 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Faster r-cnn: Towards real-time object detection with region proposal networks,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.439282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.647447Z digest=sha256:ada8bb0bf73264ce28a07e86b9c550036e43eaca5f2182b90f208d2c2ad7fa5f

Observation 820545a0-8f94-4934-b661-03f81fcd31c3 · outbound

This paper cites Stacked cross attention for image-text matching,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Stacked cross attention for image-text matching,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.421214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.652953Z digest=sha256:119de43b296ac90c10a633edab1ef94ac71d4632519c83b047e3a55ecb4d9a9e

Observation cdfd30ec-1d79-4a7c-be7b-f4ca4a409cfe · outbound

This paper cites Image- text embedding learning via visual and textual semantic reasoning,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Image- text embedding learning via visual and textual semantic reasoning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.397910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.657786Z digest=sha256:5fa8da8c4079abb1db2cc0ef4bcdb22307251b62536d7249385d59a87ed35f06

Observation 514a9348-726f-42c5-b9a6-7cd810702ce2 · outbound

This paper cites At- tentive mask clip,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching At- tentive mask clip,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.367248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.663498Z digest=sha256:55dc750bffc6d5ed53f229aab75f36de5f78eaf8fb2191d009e0e2420f467f3b

Observation c5fe10b2-5c13-4113-89f0-2502927467bb · outbound

This paper cites Composing object relations and attributes for image-text matching,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Composing object relations and attributes for image-text matching,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.352015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.669323Z digest=sha256:927dc0e6ecbdb32ba78c60ed1a596f8e71cb471a038c9634ab0a893dfdb4086c

Observation 8d94c9a9-67e6-475f-adf3-a00f85950f4e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Learning transferable visual models from natural language supervi- sion,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.336427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.675803Z digest=sha256:a66a1a1dfe328d6411a6d5d6992779e6ebf4efec4af34438794f8c97a031c5bb

Observation 8ca9b18e-f722-4b92-b8a7-a6d85806f3d2 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Sigmoid loss for language image pre-training,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.319859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.683871Z digest=sha256:7e97687e60ad0cfac3d96b57f1a26d5f7927425896adefb5675e0edd115c7bcb

Observation 2f68182e-5c82-4a45-bea4-6ba588a05ed2 · outbound

This paper cites Regionclip: Region-based language-image pretraining,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Regionclip: Region-based language-image pretraining,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.292337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.693360Z digest=sha256:c586eaafaef63adca0ff1b0f9801fca3f7461efd4733a15dc72514d2224e6154

Observation 30654223-7950-4359-9836-f79e4b7a9526 · outbound

This paper cites Sclip: Rethinking self-attention for dense vision-language inference,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Sclip: Rethinking self-attention for dense vision-language inference,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.267034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.703590Z digest=sha256:e8a48dc188aa7b021842e63b696e0370fd2d14ca375c1aae44af44f4fc6aa0b3

Observation 9bccf390-e917-45ae-b837-fc8e5ac865aa · outbound

This paper cites Improving clip training with language rewrites,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Improving clip training with language rewrites,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.250030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.715883Z digest=sha256:18b9a1bf83cc94c39313cf4645754af4e9ece5ac1154b21189af8ca0797ee784

Observation f1acdda5-029f-4367-ad21-f8a58209a01a · outbound

This paper cites Improving multimodal datasets with image captioning,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Improving multimodal datasets with image captioning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.234196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.728908Z digest=sha256:db36f49e3d50001ea83f64b55ae9a0fd70adf1ca011960bb314c80a197f249bd

Observation cbb654d8-1fa7-4e50-b034-6b2ba87c5c8a · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.218150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.735122Z digest=sha256:2ffba970061f53b72adba852831eb6ee3a74afe9d9db77f310f345991882ef36

Observation 0ee958d2-bfaf-42c2-be02-df1ebbc15013 · outbound

This paper cites MLLMs-Augmented Visual-Language Representation Learning.

Visual Semantic Description Generation with MLLMs for Image-Text Matching MLLMs-Augmented Visual-Language Representation Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.740424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.740424Z digest=sha256:3b673fce440cebef4a86198c46de0224054a0f9c236fb526fa31722be78bcf4c

Observation 6af04eb6-4472-4714-bb0f-7cc3062c3059 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Visual Semantic Description Generation with MLLMs for Image-Text Matching MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.748432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.748432Z digest=sha256:9eb75e4856e7f41cffebb4ea24f7b52634d91492690912557b2e9a1acd08853c

Observation 0033427b-733e-4acb-b098-13d0f736b924 · outbound

This paper cites C-Pack: Packed Resources For General Chinese Embeddings.

Visual Semantic Description Generation with MLLMs for Image-Text Matching C-Pack: Packed Resources For General Chinese Embeddings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.753670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.753670Z digest=sha256:287a93974d29050769ac03d6b8d8d52252f849f33064f5ab5f4986e82e1b2447

Observation 08f96dca-8f4c-404a-be87-e9c1a9a24f75 · outbound

This paper cites Fine-grained image- text matching by cross-modal hard aligning network,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Fine-grained image- text matching by cross-modal hard aligning network,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.201309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.761302Z digest=sha256:ae59b9a84f1af88a00bc3f98d9856ee3f91911d8375ed8b2b5f54926bbfb66fb

Observation 50e81f65-4615-40ad-b95f-ed89de7a1fc4 · outbound

This paper cites Fast, accurate, and lightweight memory-enhanced embedding learning framework for image-text retrieval,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Fast, accurate, and lightweight memory-enhanced embedding learning framework for image-text retrieval,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.183805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.768032Z digest=sha256:fb8da6d28e6094abf937f002c76322921599afb20376d806a1bad393dd01e483

Observation fcd0442a-4ab7-452f-a98d-8230c52abf65 · outbound

This paper cites Learning the best pooling strategy for visual semantic embedding,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Learning the best pooling strategy for visual semantic embedding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.166182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.772689Z digest=sha256:318bc9980f5507e4c292d10e88a5580961e7cae300213ae6ff6c871784d842d9

Observation e001793f-3add-4bdf-9b6a-1e4239ba249c · outbound

This paper cites Learning semantic relationship among instances for image-text matching,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Learning semantic relationship among instances for image-text matching,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.145425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.777092Z digest=sha256:cb21483311d293bb27ddc4cda25f766ad0c87c3b33f734fa213090dc772fffcf

Observation 26a67493-2443-427f-b1e0-30aa4136f551 · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Sinkhorn distances: Lightspeed computation of optimal transport,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.124932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.797389Z digest=sha256:d02970cd66f7dd1d3f56d7d6c0dcb14bc461915774acb257260c24686c5f5dca

Observation f2d0d17a-2ff4-4404-9b24-cbaf2c7c7a03 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Deep visual-semantic alignments for generating image descriptions,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.105418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.805931Z digest=sha256:215f9976cc8c721ff635ffcc195af5ab24cecca294cf6e2696b112431fc4dd1b

Observation 593e71a4-bd87-4236-a740-ca1389172458 · outbound

This paper cites N24News: A New Dataset for Multimodal News Classification.

Visual Semantic Description Generation with MLLMs for Image-Text Matching N24News: A New Dataset for Multimodal News Classification

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.811615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.811615Z digest=sha256:2f0e6e0215f90146693032728b270fee3035a6c903a0a3a1fd29c8ac3c2541d8

Observation e83deb29-8974-4017-8f07-bfd8cee1eefa · outbound

This paper cites Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.084295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.817342Z digest=sha256:3929f645ebbc400ec815c34c0201bcb75bfd49024c71a1ef16517c9459daaebe

Observation b76fc5ee-b355-4008-8d7e-2ea5b2803cc3 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Florence-2: Advancing a unified representation for a variety of vision tasks,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.065552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.825844Z digest=sha256:3ca6df92249495cd36d8c462305771c3a9383389add6860b647166e72ce57f18

Observation 337190bb-dc06-4ccb-9f17-3e968b6363a9 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Align before fuse: Vision and language representation learning with momentum distillation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.045119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.832266Z digest=sha256:d3ffdfb76180ee091e1d34fdb12fb37137783671cbe559a063edc59f950dd321

Observation 8ab328d0-2d7d-4a6b-bbe6-4b9ed6e2c7ca · outbound

This paper cites Remote sensing cross-modal text-image retrieval based on global and local information,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Remote sensing cross-modal text-image retrieval based on global and local information,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.019781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.842725Z digest=sha256:53ec0e953e8afe098d4f457f1925f792cd70d2a80ade4f34aabe2317b78c94fc

Observation 5d951393-c6d2-46c2-b4a3-189abee9d957 · outbound

This paper cites Hypersphere-based remote sensing cross-modal text-image retrieval via curriculum learning,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Hypersphere-based remote sensing cross-modal text-image retrieval via curriculum learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.000703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:19:58.847247Z digest=sha256:7a994ed17f409364902659b9dc23ea5b4a6ec2fe309b2b77743db1aafee63747

Pith citing papers

No inbound Pith citation observations are available.