Pith. sign in

Paper Citation Record · LEDGER

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation

As of 19 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2412.10761.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10761 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:42:28.284177Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1239b7d4-e612-4627-8a0b-131fae79a59b · outbound

This paper cites Vilbert: Pre- training task-agnostic visiolinguistic representations for vision-and-language tasks,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Vilbert: Pre- training task-agnostic visiolinguistic representations for vision-and-language tasks,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.248782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:27.999890Z digest=sha256:5935571019e79332ae34a0f7a86118a907a6fc5de3159b294d6f7310d46b00e9

Observation adfb622f-6abc-4bde-adf6-4545f45d3e92 · outbound

This paper cites Explore instance similarity: An in- stance correlation based hashing method for multi-label cross-model retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Explore instance similarity: An in- stance correlation based hashing method for multi-label cross-model retrieval,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.233043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.005941Z digest=sha256:f40c8f0b5b5bcf6cc21f40ce28ab708712e2b295cf780518cca4d13497190935

Observation 57acdcb3-d53e-420a-8311-d5d55dd634d1 · outbound

This paper cites Dynamic contrastive distillation for image-text retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Dynamic contrastive distillation for image-text retrieval,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.216828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.011094Z digest=sha256:e090e27b7f119524f12f435dc01038e206cf1fcd4d22cf3d86ee640189904043

Observation 35fd993e-a450-4ffb-99bd-867721edd71d · outbound

This paper cites Covlr: Coordi- nating cross-modal consistency and intra-modal relations for vision-language retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Covlr: Coordi- nating cross-modal consistency and intra-modal relations for vision-language retrieval,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.200935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.016219Z digest=sha256:66a8c2f450c063a9c2c55b7757f2c1af084b726fe046ee85b5bab66a6fd496e7

Observation 354e6e41-1021-4e05-8af5-e3a594e3e976 · outbound

This paper cites Deep relation embedding for cross-modal retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Deep relation embedding for cross-modal retrieval,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.184738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.021170Z digest=sha256:7b9eeb5265d00391a6981da61aa888e29ef3706c23548f4c85e3fb9bb6afa01d

Observation 50c8c704-7147-4030-9fe2-e98ef76ff70e · outbound

This paper cites Stacked cross attention for image-text matching,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Stacked cross attention for image-text matching,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.168613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.025802Z digest=sha256:5e985a4c5c1c702cf452de5e69d2041f4c2b81f87751e1937c642276e71b5197

Observation e9fb61ca-145a-4f7e-a883-840ad84ab2c5 · outbound

This paper cites Cross- modal retrieval with partially mismatched pairs,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Cross- modal retrieval with partially mismatched pairs,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.152702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.031672Z digest=sha256:a585746a3ee716d0fd3f15cd618a241d90e6e0263d10ffb9e309b4558b80a683

Observation d714d033-7122-4f58-9990-ae090dd0e9f4 · outbound

This paper cites VSE++: improving visual-semantic embeddings with hard nega- tives,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation VSE++: improving visual-semantic embeddings with hard nega- tives,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.136511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.036418Z digest=sha256:88238140812663bfbf13e0197f02eabf91c75fdebc892d0ce7e996d1ea9f7403

Observation c935032c-ba13-45a9-b7f2-75bd2eb337c0 · outbound

This paper cites Modality-specific cross- modal similarity measurement with recurrent attention network,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Modality-specific cross- modal similarity measurement with recurrent attention network,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.104002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.046002Z digest=sha256:a01c0b43f54befe261de955f5be4b59879804940328f526877465c6d22893109

Observation 521bf46d-763b-4c28-9736-5616af8ad40a · outbound

This paper cites Show your faith: Cross-modal confidence-aware network for image- text matching,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Show your faith: Cross-modal confidence-aware network for image- text matching,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.086748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.050976Z digest=sha256:8d4fd31f9c054a3e9370208898106e5cb64dd112870a336eba4773b79d60f35f

Observation 35e6e395-2ccb-481b-82a0-409c2449c9a6 · outbound

This paper cites Similarity rea- soning and filtration for image-text matching,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Similarity rea- soning and filtration for image-text matching,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.120410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.056043Z digest=sha256:09508a9cf7f8dfef2d7d819bfe9f3d4bf5dd60e850afdd030c1d2bca6c63c02a

Observation da98783f-b48d-46ea-b59a-d2d1e0fbf280 · outbound

This paper cites Towards lightweight transformer IEEE TRANSACTIONS ON IMAGE PROCESSING, DECEMBER 2024 12 via group-wise transformation for vision-and-language tasks,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Towards lightweight transformer IEEE TRANSACTIONS ON IMAGE PROCESSING, DECEMBER 2024 12 via group-wise transformation for vision-and-language tasks,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.070084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.060979Z digest=sha256:5c026a320c2185f400ff061182f10ed0c23b43cd54081e523f76a5f5587bd476

Observation dacc8a09-ca20-46eb-8eaa-368d7cd23813 · outbound

This paper cites BLIP: boot- strapping language-image pre-training for unified vision- language understanding and generation,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation BLIP: boot- strapping language-image pre-training for unified vision- language understanding and generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.053248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.065899Z digest=sha256:d45b1aeb981322c30618a1f385492dae5eeb2be42f99f32ce58b57c612cd50a1

Observation 79485594-ffc3-4edc-a93b-e0d3559e4b27 · outbound

This paper cites Learning semantic relationship among instances for image-text matching,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Learning semantic relationship among instances for image-text matching,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.036038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.070714Z digest=sha256:4ab9ba1b36c0b66176f4a83120e1b1afa40b8112afcecbd4e35ee52c0cbf54f9

Observation bb9bfdcf-36fc-445a-8c39-a02dab175cd7 · outbound

This paper cites Co-training with insufficient views,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Co-training with insufficient views,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.018578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.075764Z digest=sha256:ca770787226b35d56e1d380fbb0d7f3c883ae80b8e669e1c9214bf83d365e80a

Observation e6bebe61-0add-4e05-b761-302e18295437 · outbound

This paper cites Multi-modal mutual attention and iterative interaction for referring image segmentation,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Multi-modal mutual attention and iterative interaction for referring image segmentation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.999862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.080609Z digest=sha256:41755e140424dd39cb7b88a76888060dc20e2514cbddbd4cf67c16b654871993

Observation 3018a639-0b03-4166-a39b-7516f07a50f6 · outbound

This paper cites Auxiliary information regularized machine for multiple modality feature learning,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Auxiliary information regularized machine for multiple modality feature learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.983479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.085742Z digest=sha256:8ac29aa078e179c946f8ade4cb4402fe8816f1349cca23add91e796e48444acb

Observation e6029caa-7ccd-424f-94c6-cbc9ef82f26f · outbound

This paper cites Modality competition: What makes joint training of multi-modal network fail in deep learning? (provably),.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Modality competition: What makes joint training of multi-modal network fail in deep learning? (provably),

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.966623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.090585Z digest=sha256:eb01915f9a461512c70fa13b75bfa89f9f44fc4469c4b24235ea83b3f214b34a

Observation 578d0eb8-cff5-42d4-8beb-4d1c8fcd2a8e · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for lan- guage understanding,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation BERT: pre-training of deep bidirectional transformers for lan- guage understanding,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.949795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.095497Z digest=sha256:04b02bfd2ff078b0dc2bc70e5c0e5ab009d53822c2aef665e50f9a68bf7637dd

Observation 36720ff9-8ca1-43ae-8512-259f5f25036b · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Swin transformer: Hierarchical vision transformer using shifted windows,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.933138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.100278Z digest=sha256:955de28182ff088ae3b37c28344c4eaf7ffadef22be9de17b07b5d1352ef1007

Observation 47273d1f-a8db-4662-b99f-cd94c39521dd · outbound

This paper cites Visualizing data using t-sne.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Visualizing data using t-sne

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:28.104943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:28.104943Z digest=sha256:cd5aebdf2398d9191d34f73559ea95fa555c61aa79fcd18ec769889ac6f11919

Observation bde5093d-13af-4558-8ef1-bf9b1302c991 · outbound

This paper cites Rethinking label-wise cross-modal retrieval from A semantic sharing perspective,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Rethinking label-wise cross-modal retrieval from A semantic sharing perspective,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.902668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.109529Z digest=sha256:9d17e6c299b8741422ec08728acfb58f84b83d47ef169e0ea872f6f0dd23e521

Observation 051ea1b3-817b-40a2-813b-3b3e70bb8e48 · outbound

This paper cites an unresolved cited work.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:42:28.885981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.114406Z digest=sha256:34b472d3439ad2964908c19ddb4506629ea6e945eb271516b7cbe88d4c29d79e

Observation bc50d7c9-aeff-4e7a-b67e-3303ab6de79f · outbound

This paper cites Joint feature synthesis and embedding: Adversarial cross-modal retrieval revisited,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Joint feature synthesis and embedding: Adversarial cross-modal retrieval revisited,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.869062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.119152Z digest=sha256:8187170090456a3187729e1976ebd70e37fe80afeb74a5b2961ad1f62c3a9554

Observation a33c4c2f-5235-4f92-83c6-ed9afb76c422 · outbound

This paper cites Ad- versarial graph convolutional network for cross-modal retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Ad- versarial graph convolutional network for cross-modal retrieval,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.852176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.123955Z digest=sha256:16b735396827d92d18f0b6232f035e8ed5598bf862a51a51065724ad666b04f5

Observation 93c78f73-509a-4bd5-848b-30567a467c4a · outbound

This paper cites Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.836046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.128806Z digest=sha256:fdba196c94708f7e9276fca99f1a76c2de395b207d1bbedd16f44be4e7ee993f

Observation 97021fae-b162-46c2-99ec-70bbb91a47b2 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Learning transferable visual models from natural language supervision,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:28.134205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:28.134205Z digest=sha256:3a8411fc657faab4915ab55a8b8b5ed773ec0ae3ec501e860cb5136021e37607

Observation cf41082e-8a5f-4621-a86c-2b29a127e332 · outbound

This paper cites Unsupervised contrastive cross-modal hashing,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Unsupervised contrastive cross-modal hashing,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.808444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.139094Z digest=sha256:e73f58fbc1a7aebcaa6a3ece6f03a67820dc61a901ee768cf8077fb66eabf3b9

Observation 88ed5a55-c026-4b5e-9523-4722632eef91 · outbound

This paper cites LXMERT: learning cross- modality encoder representations from transformers,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation LXMERT: learning cross- modality encoder representations from transformers,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.792708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.143999Z digest=sha256:ad71979f210b52e54dee5c527db170eaac7b8357e498819c5cea03fd284ff372

Observation 1ccaad70-a3c8-4730-9477-9dfc598a8436 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Align before fuse: Vision and language representation learning with momentum distillation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.776814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.148752Z digest=sha256:0cf7d46a264b3a470eec1024de6a0dc9cf8b80e584505e34bfaa9a186de0d741

Observation b20a4390-6347-48ac-8a43-f2482e59c271 · outbound

This paper cites What makes training multi-modal classification networks hard?.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation What makes training multi-modal classification networks hard?

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.761112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.153567Z digest=sha256:e2bafc28b2490afab0d22e460c4627c808337b0405184fe992733f6fa1afec35

Observation 4b0d9c84-77d3-42be-8556-c429ec79e0d8 · outbound

This paper cites Trusted multi- view classification,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Trusted multi- view classification,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.744935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.158656Z digest=sha256:06e3d6e50abea042def6bc38a26f3484cbe10335eddb5367e900e428ccde3f8c

Observation f70341bf-3e9b-4d99-af95-8c0843083792 · outbound

This paper cites Balanced multimodal learning via on-the-fly gradient modulation,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Balanced multimodal learning via on-the-fly gradient modulation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.729218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.163526Z digest=sha256:2139df7890a7472ef0cf8e5d24af07f40c2dbc58b8b45d4a1bb7376fa408b159

Observation bd5eda0d-a85c-4517-a671-68e17c3be8d9 · outbound

This paper cites Multi-grained vision lan- guage pre-training: Aligning texts with visual concepts,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Multi-grained vision lan- guage pre-training: Aligning texts with visual concepts,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.713139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.168411Z digest=sha256:179a7c8c5b55b80de6f643b0f99afa530e222291726e3e8e2682839e89a7dc84

Observation 3d59aadf-2b35-46fb-a452-dce10608da94 · outbound

This paper cites Deep residual learning for image recognition,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Deep residual learning for image recognition,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.697114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.173683Z digest=sha256:9e931f2ed70c4bc05bd795f77bb4b2138686059bacea67d49c4199f3ba08e52d

Observation 4f39dd36-ed2e-4d3a-8f55-65e86da9f12d · outbound

This paper cites Long short-term memory,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Long short-term memory,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.681814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.178597Z digest=sha256:244fe774a4fd56ebb8f72f39e6f2c6fcdaa2337e9ba547b3ab593e3dd6e84f5f

Observation c84b2b5a-605b-4434-8c51-475bc4fd3d9d · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.666145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.183595Z digest=sha256:bc418cbbbbaaa59a01c9c1f06fcfe00a55980368e7e9d145dfefae7c04425c9c

Observation f4d5647d-7812-459d-8a14-9dbd5b946213 · outbound

This paper cites Relational knowledge distillation,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Relational knowledge distillation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.649954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.188307Z digest=sha256:bfe7175e0df0f5457f7ca53cb0ef26351bd5a994dfcb6b255cd4afbf6135063f

Observation 6482f933-c931-4496-baf2-5c3b81ab1212 · outbound

This paper cites Large-margin contrastive learning with distance polarization regularizer,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Large-margin contrastive learning with distance polarization regularizer,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.632927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.193064Z digest=sha256:08209795edf7c647dab399a256c77e524bdbf5cac4ef669e3f324b915ce6706a

Observation 9d8f6a96-c8bb-4e90-8920-0fae10a8618b · outbound

This paper cites DINO: DETR with improved denoising anchor boxes for end-to-end object detection,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation DINO: DETR with improved denoising anchor boxes for end-to-end object detection,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.617115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.197656Z digest=sha256:dbdddfe9f78d06ccc9478e8780d4bbedc83dbb9d5d4e07284d58c7f0a4c04a59

Observation dc1d3cab-9fef-4c92-a1fa-8b83433e8ee2 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.597980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.202561Z digest=sha256:e1dfa93bdc081da5be786051893c5376a7784dfbc974af0513646b1a8f4fb8b6

Observation c3d1ff54-9a73-46fe-bcae-e60fa146a093 · outbound

This paper cites Microsoft coco: Common objects in context,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Microsoft coco: Common objects in context,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.579938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.207517Z digest=sha256:8895ce985377026631a5b55f49b01d8668eaecf9227b05946c8fbdeafdaffe6e

Observation 634b4e69-4e92-4d94-82a7-c7d54b234a20 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Deep visual-semantic alignments for generating image descriptions,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.563748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.212825Z digest=sha256:926776578405cd46cb662824ec6c8571d4ee8f8f1b55701f0e0a73ebb1293165

Observation 1d472949-ad97-4099-8262-c19cafc6823e · outbound

This paper cites The MIR flickr retrieval evaluation,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation The MIR flickr retrieval evaluation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.547232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.217660Z digest=sha256:e8331b528bc76f44aaeb4d7e986d3f1872f1431b7f901b523b87e5e26da18318

Observation 43b840e8-ff88-4dc0-a380-ee49f7372bc2 · outbound

This paper cites Captioning images taken by people who are blind,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Captioning images taken by people who are blind,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.530673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.222527Z digest=sha256:387d453d902ba38d9c48826c5b1bfe02de9164f11a2decc1974f24c03423122d

Observation eaa6dfd3-e219-497d-8468-99e636f1ef66 · outbound

This paper cites IMRAM: iterative matching with recurrent attention memory for cross-modal image-text retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation IMRAM: iterative matching with recurrent attention memory for cross-modal image-text retrieval,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.513876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.227467Z digest=sha256:34b5bca039198c4974f6ada59e0bf7ea5d6d69a4cc06e529fd91a2af50101117

Observation d6ce78ee-f3a6-41d4-98e7-9c2c0be2a5cf · outbound

This paper cites Graph structured network for image-text matching,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Graph structured network for image-text matching,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.497055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.232464Z digest=sha256:6fb88442f1147b39b65451f49167361c8823ad1f48585ea5fb82f02372e49836

Observation f7faea77-7c8d-4586-9baf-aa02ab69adfc · outbound

This paper cites Visual semantic reasoning for image-text matching,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Visual semantic reasoning for image-text matching,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.480530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.237334Z digest=sha256:a49a2280430378db75b46237c4ec470165eba93d69a32762e4559c2effc9446d

Observation 2855ac7c-89c5-4e36-9ca9-fb8c6aad4fa7 · outbound

This paper cites Negative- aware attention framework for image-text matching,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Negative- aware attention framework for image-text matching,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.464004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.242506Z digest=sha256:314639608c64a55a583612ccd063492dbdab056fb1b88dbc72bdce8f266edf5f

Observation 0c59de8c-3129-4e56-ab57-d752aae450ae · outbound

This paper cites Cyclip: Cyclic contrastive language-image pretraining,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Cyclip: Cyclic contrastive language-image pretraining,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.447654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.247977Z digest=sha256:cce21bc6e1933cc8be1bd1585adf3502557a5fac168219567ee2e1ec9ff39527

Observation 0b6c0d14-0423-41e7-bff8-8fa542e372f8 · outbound

This paper cites Improving Multi-Modal Learning with Uni-Modal Teachers.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Improving Multi-Modal Learning with Uni-Modal Teachers

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:28.253438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:28.253438Z digest=sha256:60ad871cc8dcbc46160a77d24c35e14786c715c74f89b6b0a131580508c13c12

Observation 754cb223-4ff8-49d4-a972-b414c9b0019d · outbound

This paper cites Neighborhood discriminant hashing for large-scale image retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Neighborhood discriminant hashing for large-scale image retrieval,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.430514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.258989Z digest=sha256:70de30d97cb07ef8d6e80201452315f234d4427f60dd0475680b05791b403e4e

Observation f701a0fe-790c-402e-8af7-22e42d96a677 · outbound

This paper cites Learning dis- criminative cross-modality features for rgb-d saliency detection,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Learning dis- criminative cross-modality features for rgb-d saliency detection,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.412090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.264147Z digest=sha256:f19aecb22ebb3ec3a7ba0bdee1cb78aed4aa18fdd2013b486e36b8d76a42dabb

Observation 04bc54b9-f5b4-499b-99a7-cc263718e732 · outbound

This paper cites Camera constraint-free view-based 3-d object retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Camera constraint-free view-based 3-d object retrieval,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.395834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.268958Z digest=sha256:5c3156432c541f69334f12b2dd7c91dd14e06b69bf3662e7bdc60065092bcfad

Observation 124c95db-928b-4e4b-a04b-b7521a1177eb · outbound

This paper cites Picture it in your mind: generating high- level visual representations from textual descriptions,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Picture it in your mind: generating high- level visual representations from textual descriptions,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.379885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.273983Z digest=sha256:b0987360bae715cf098aa06350a94bf1db5649425f959e1aad4d7b7adcad36b7

Observation 104a5a55-ba63-401e-8b87-808f7872fa42 · outbound

This paper cites Attention on attention for image captioning,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Attention on attention for image captioning,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.362920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.279004Z digest=sha256:5bd3fde56643f636f130da19187e18760de8a973b437924306db66e7f9d7c434

Observation bdac5ead-e5a5-4b5c-bab6-46a6d67db623 · outbound

This paper cites Decoupled weight decay regularization,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Decoupled weight decay regularization,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.345018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T15:42:28.284177Z digest=sha256:784ce90c4e2a4907aec650e07b4d588ab3059cdc7103a5b450227715747ba784

Pith citing papers

No inbound Pith citation observations are available.