Pith. sign in

Paper Citation Record · LEDGER

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation

As of 17 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2412.10761.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10761 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:42:28.284177Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1239b7d4-e612-4627-8a0b-131fae79a59b · outbound

This paper cites Vilbert: Pre- training task-agnostic visiolinguistic representations for vision-and-language tasks,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Vilbert: Pre- training task-agnostic visiolinguistic representations for vision-and-language tasks,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.248782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:27.999890Z digest=sha256:bd11ec63b76c74974d8e89becf18c210b95931e976deb0e62c7c5d4aabeab7af

Observation adfb622f-6abc-4bde-adf6-4545f45d3e92 · outbound

This paper cites Explore instance similarity: An in- stance correlation based hashing method for multi-label cross-model retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Explore instance similarity: An in- stance correlation based hashing method for multi-label cross-model retrieval,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.233043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.005941Z digest=sha256:3193b7b8bf8d2f10fc497fd9ab269b8316ff64b7d56a140846193b703e01ba16

Observation 57acdcb3-d53e-420a-8311-d5d55dd634d1 · outbound

This paper cites Dynamic contrastive distillation for image-text retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Dynamic contrastive distillation for image-text retrieval,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.216828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.011094Z digest=sha256:7b1e5459e237fb41d84a6976acd5e6987daf612b0fbd007e1720f9ab885339af

Observation 35fd993e-a450-4ffb-99bd-867721edd71d · outbound

This paper cites Covlr: Coordi- nating cross-modal consistency and intra-modal relations for vision-language retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Covlr: Coordi- nating cross-modal consistency and intra-modal relations for vision-language retrieval,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.200935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.016219Z digest=sha256:5d98a5e5cf6d3620e4a089a075486d9ae1d11922e53264f14b1968d28450ebb9

Observation 354e6e41-1021-4e05-8af5-e3a594e3e976 · outbound

This paper cites Deep relation embedding for cross-modal retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Deep relation embedding for cross-modal retrieval,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.184738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.021170Z digest=sha256:7fdb066d6a8583929a2a43a18eb63ed68d57de2397e85634bba3fa2676e7cef6

Observation 50c8c704-7147-4030-9fe2-e98ef76ff70e · outbound

This paper cites Stacked cross attention for image-text matching,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Stacked cross attention for image-text matching,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.168613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.025802Z digest=sha256:0435ec714105ddd059a1a84eac41cdeea2a24aa1735897205aed412946a45144

Observation e9fb61ca-145a-4f7e-a883-840ad84ab2c5 · outbound

This paper cites Cross- modal retrieval with partially mismatched pairs,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Cross- modal retrieval with partially mismatched pairs,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.152702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.031672Z digest=sha256:4fc38328b46bc573b4c838add22be2f11bb2e354c16a29ba64a64cbde8d05e8b

Observation d714d033-7122-4f58-9990-ae090dd0e9f4 · outbound

This paper cites VSE++: improving visual-semantic embeddings with hard nega- tives,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation VSE++: improving visual-semantic embeddings with hard nega- tives,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.136511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.036418Z digest=sha256:82ad010465d332297a4a78dfc4f911b2800d02efbc0edeae37289b53485a0a15

Observation c935032c-ba13-45a9-b7f2-75bd2eb337c0 · outbound

This paper cites Modality-specific cross- modal similarity measurement with recurrent attention network,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Modality-specific cross- modal similarity measurement with recurrent attention network,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.104002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.046002Z digest=sha256:669274ee9c50c3d81dc9d5b5af3839b2e22921f2f63f89d31e4181c7d1a42357

Observation 521bf46d-763b-4c28-9736-5616af8ad40a · outbound

This paper cites Show your faith: Cross-modal confidence-aware network for image- text matching,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Show your faith: Cross-modal confidence-aware network for image- text matching,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.086748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.050976Z digest=sha256:782762e1b9b5d0f5e2e4594ee926ff93a3c648be471ea456c22e924ee84c8add

Observation 35e6e395-2ccb-481b-82a0-409c2449c9a6 · outbound

This paper cites Similarity rea- soning and filtration for image-text matching,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Similarity rea- soning and filtration for image-text matching,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.120410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.056043Z digest=sha256:0724826a63c5123b52fe1dd6773c3dbf7b7c2ceed7a3f44330656aae2dc330b4

Observation da98783f-b48d-46ea-b59a-d2d1e0fbf280 · outbound

This paper cites Towards lightweight transformer IEEE TRANSACTIONS ON IMAGE PROCESSING, DECEMBER 2024 12 via group-wise transformation for vision-and-language tasks,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Towards lightweight transformer IEEE TRANSACTIONS ON IMAGE PROCESSING, DECEMBER 2024 12 via group-wise transformation for vision-and-language tasks,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.070084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.060979Z digest=sha256:00452ba1f9066580ae9751ab00c09e20ca20631af7cbf59ca47a528bf6bda7c6

Observation dacc8a09-ca20-46eb-8eaa-368d7cd23813 · outbound

This paper cites BLIP: boot- strapping language-image pre-training for unified vision- language understanding and generation,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation BLIP: boot- strapping language-image pre-training for unified vision- language understanding and generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.053248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.065899Z digest=sha256:2990d4321b28bcb28f8ca43e0817f82c2436b1d5ad5b69b1eefa3fa09e5928ea

Observation 79485594-ffc3-4edc-a93b-e0d3559e4b27 · outbound

This paper cites Learning semantic relationship among instances for image-text matching,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Learning semantic relationship among instances for image-text matching,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.036038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.070714Z digest=sha256:5d378d0e44eb7d448abc2654df91a2e653b06ad26b2984b7f5e28b0cfad40b19

Observation bb9bfdcf-36fc-445a-8c39-a02dab175cd7 · outbound

This paper cites Co-training with insufficient views,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Co-training with insufficient views,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:29.018578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.075764Z digest=sha256:10232d5115d5af5a85125df1690a95928ed9f78a40e9a36a73dc26a01e558866

Observation e6bebe61-0add-4e05-b761-302e18295437 · outbound

This paper cites Multi-modal mutual attention and iterative interaction for referring image segmentation,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Multi-modal mutual attention and iterative interaction for referring image segmentation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.999862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.080609Z digest=sha256:fafd7cccdcb582660275e5dae99684734fc131ff58c61d74a1caea3b469f3f3c

Observation 3018a639-0b03-4166-a39b-7516f07a50f6 · outbound

This paper cites Auxiliary information regularized machine for multiple modality feature learning,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Auxiliary information regularized machine for multiple modality feature learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.983479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.085742Z digest=sha256:0f821c05c76d476b6baa269b342d429f948184161d343fda9225ecd044e79023

Observation e6029caa-7ccd-424f-94c6-cbc9ef82f26f · outbound

This paper cites Modality competition: What makes joint training of multi-modal network fail in deep learning? (provably),.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Modality competition: What makes joint training of multi-modal network fail in deep learning? (provably),

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.966623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.090585Z digest=sha256:0b103204f6feb342fdde5dc2fd127ff4911798c78abf79c7180dd8eb197a25e7

Observation 578d0eb8-cff5-42d4-8beb-4d1c8fcd2a8e · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for lan- guage understanding,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation BERT: pre-training of deep bidirectional transformers for lan- guage understanding,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.949795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.095497Z digest=sha256:7a38123afc67905b6bded4e21aedfd518378428e453c90634278d2c150c9cffe

Observation 36720ff9-8ca1-43ae-8512-259f5f25036b · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Swin transformer: Hierarchical vision transformer using shifted windows,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.933138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.100278Z digest=sha256:59ac925c3b0362abd1eef5fa0c7260b0ab9701b318c77a55bb7e1173c9ba4e65

Observation 47273d1f-a8db-4662-b99f-cd94c39521dd · outbound

This paper cites Visualizing data using t-sne.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Visualizing data using t-sne

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:28.104943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:28.104943Z digest=sha256:cd5aebdf2398d9191d34f73559ea95fa555c61aa79fcd18ec769889ac6f11919

Observation bde5093d-13af-4558-8ef1-bf9b1302c991 · outbound

This paper cites Rethinking label-wise cross-modal retrieval from A semantic sharing perspective,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Rethinking label-wise cross-modal retrieval from A semantic sharing perspective,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.902668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.109529Z digest=sha256:105ed608e926be88e6c94dbcf7c83d5f35bdcda3265c0ea5d55a3f744744c46e

Observation 051ea1b3-817b-40a2-813b-3b3e70bb8e48 · outbound

This paper cites an unresolved cited work.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:42:28.885981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.114406Z digest=sha256:39e00c69888216fdb9ae42a01119cc0e65c89a49228ce9ca649f1faf2d7cf91c

Observation bc50d7c9-aeff-4e7a-b67e-3303ab6de79f · outbound

This paper cites Joint feature synthesis and embedding: Adversarial cross-modal retrieval revisited,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Joint feature synthesis and embedding: Adversarial cross-modal retrieval revisited,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.869062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.119152Z digest=sha256:ccd8b869344e28b1aaccc14b5cd6fd365f832cdbf634da2f25f7323b761ade42

Observation a33c4c2f-5235-4f92-83c6-ed9afb76c422 · outbound

This paper cites Ad- versarial graph convolutional network for cross-modal retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Ad- versarial graph convolutional network for cross-modal retrieval,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.852176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.123955Z digest=sha256:74657a22bb812c4d162076881e1bd0faac20af9bbafad24122dfb229087dea78

Observation 93c78f73-509a-4bd5-848b-30567a467c4a · outbound

This paper cites Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Fine-grained visual textual alignment for cross-modal retrieval using transformer encoders,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.836046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.128806Z digest=sha256:01ddfaffc400b0c1ed0036a88bc1dec08179beeff673181fc5da69f6dad85f31

Observation 97021fae-b162-46c2-99ec-70bbb91a47b2 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Learning transferable visual models from natural language supervision,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:28.134205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:28.134205Z digest=sha256:3a8411fc657faab4915ab55a8b8b5ed773ec0ae3ec501e860cb5136021e37607

Observation cf41082e-8a5f-4621-a86c-2b29a127e332 · outbound

This paper cites Unsupervised contrastive cross-modal hashing,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Unsupervised contrastive cross-modal hashing,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.808444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.139094Z digest=sha256:43563c8c85602e03392b96da7e7c713f226217d227ab4896075e5eb83c24c68f

Observation 88ed5a55-c026-4b5e-9523-4722632eef91 · outbound

This paper cites LXMERT: learning cross- modality encoder representations from transformers,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation LXMERT: learning cross- modality encoder representations from transformers,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.792708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.143999Z digest=sha256:3a02b71d9b003a0435dd840b5754c09fb3b13782747fdea9c4002258692d4e1d

Observation 1ccaad70-a3c8-4730-9477-9dfc598a8436 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Align before fuse: Vision and language representation learning with momentum distillation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.776814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.148752Z digest=sha256:e176e8b829bd462c1452b9e8e3646dc1faec7ae6d8266daffbb7def905d3adde

Observation b20a4390-6347-48ac-8a43-f2482e59c271 · outbound

This paper cites What makes training multi-modal classification networks hard?.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation What makes training multi-modal classification networks hard?

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.761112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.153567Z digest=sha256:4a905a2e74fea5245464fc736311d6fe0938cdfed3da4dd76a50b09298ab7d5a

Observation 4b0d9c84-77d3-42be-8556-c429ec79e0d8 · outbound

This paper cites Trusted multi- view classification,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Trusted multi- view classification,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.744935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.158656Z digest=sha256:4d29ebf1946db518452a0ac06f4a752f573147a7ef00a67b56822b48d198f118

Observation f70341bf-3e9b-4d99-af95-8c0843083792 · outbound

This paper cites Balanced multimodal learning via on-the-fly gradient modulation,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Balanced multimodal learning via on-the-fly gradient modulation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.729218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.163526Z digest=sha256:b5f0c09a0bbcaca5f4c96bdcb7b8e3df9ef02b00d8ae9712cdda5bcc33e3fec7

Observation bd5eda0d-a85c-4517-a671-68e17c3be8d9 · outbound

This paper cites Multi-grained vision lan- guage pre-training: Aligning texts with visual concepts,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Multi-grained vision lan- guage pre-training: Aligning texts with visual concepts,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.713139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.168411Z digest=sha256:a7efb8c413e7780a0db417d6b0164daa3b23d12655b4100ebb7a6967460d5485

Observation 3d59aadf-2b35-46fb-a452-dce10608da94 · outbound

This paper cites Deep residual learning for image recognition,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Deep residual learning for image recognition,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.697114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.173683Z digest=sha256:b028e0c953018cf3a4ef681be7d254868eab9dde1a38200c45768f3fc76d5636

Observation 4f39dd36-ed2e-4d3a-8f55-65e86da9f12d · outbound

This paper cites Long short-term memory,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Long short-term memory,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.681814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.178597Z digest=sha256:86d9963ddd35dba4a5a82adaff68abeb5631033a08682941a460007509fbee2e

Observation c84b2b5a-605b-4434-8c51-475bc4fd3d9d · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.666145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.183595Z digest=sha256:167fc505f6750184eb01249a99fd12980fe2caebfe1e295a8538a9c1cee578c9

Observation f4d5647d-7812-459d-8a14-9dbd5b946213 · outbound

This paper cites Relational knowledge distillation,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Relational knowledge distillation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.649954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.188307Z digest=sha256:3a65a67fe3c1a0a5c4ed33158552bee7719036ef823110a1dbe0bbca93f3d51c

Observation 6482f933-c931-4496-baf2-5c3b81ab1212 · outbound

This paper cites Large-margin contrastive learning with distance polarization regularizer,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Large-margin contrastive learning with distance polarization regularizer,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.632927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.193064Z digest=sha256:c2eb43536e1bc0ee682391c5fb278ab072beec3f80bcff7e49611434743091ae

Observation 9d8f6a96-c8bb-4e90-8920-0fae10a8618b · outbound

This paper cites DINO: DETR with improved denoising anchor boxes for end-to-end object detection,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation DINO: DETR with improved denoising anchor boxes for end-to-end object detection,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.617115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.197656Z digest=sha256:83c478daa171086e1c8be57eecebc4dc0380b53c4e78b7b59805b8545c8619f9

Observation dc1d3cab-9fef-4c92-a1fa-8b83433e8ee2 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.597980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.202561Z digest=sha256:c41aed14d527ed5d96acb8efde037f629177ba4cd5cdf2d60f734fba903fa0c0

Observation c3d1ff54-9a73-46fe-bcae-e60fa146a093 · outbound

This paper cites Microsoft coco: Common objects in context,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Microsoft coco: Common objects in context,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.579938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.207517Z digest=sha256:ca1801c52806556c06adad4eaeabb0a0569876ce2c186d24ee01cf9cbb944d93

Observation 634b4e69-4e92-4d94-82a7-c7d54b234a20 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Deep visual-semantic alignments for generating image descriptions,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.563748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.212825Z digest=sha256:f7452c36d6dad414d47aaa02e4809911c71984996688a7bbabf4c5231b6b43ef

Observation 1d472949-ad97-4099-8262-c19cafc6823e · outbound

This paper cites The MIR flickr retrieval evaluation,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation The MIR flickr retrieval evaluation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.547232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.217660Z digest=sha256:52665d0ba0cde16ac265782dffdacf5f921dc3e8c1c6640c8d6ad5add3457435

Observation 43b840e8-ff88-4dc0-a380-ee49f7372bc2 · outbound

This paper cites Captioning images taken by people who are blind,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Captioning images taken by people who are blind,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.530673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.222527Z digest=sha256:aff453b72bb269fb2296b8ea306edb877f707dd639f2cb4da5540315e4ade6af

Observation eaa6dfd3-e219-497d-8468-99e636f1ef66 · outbound

This paper cites IMRAM: iterative matching with recurrent attention memory for cross-modal image-text retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation IMRAM: iterative matching with recurrent attention memory for cross-modal image-text retrieval,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.513876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.227467Z digest=sha256:f69e58aaaeb4c21f2fc0b14bce65df1eceb81d38ee012bc6e5faab6cf1f66d7b

Observation d6ce78ee-f3a6-41d4-98e7-9c2c0be2a5cf · outbound

This paper cites Graph structured network for image-text matching,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Graph structured network for image-text matching,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.497055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.232464Z digest=sha256:d520d7b1090bd1508b351e0d5cc0a988a668307c4ee837b41c3113abbd05c91f

Observation f7faea77-7c8d-4586-9baf-aa02ab69adfc · outbound

This paper cites Visual semantic reasoning for image-text matching,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Visual semantic reasoning for image-text matching,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.480530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.237334Z digest=sha256:3202b78365cc4b35da46efc69023343e7d530fbf9dff583a14eb4e34227fb459

Observation 2855ac7c-89c5-4e36-9ca9-fb8c6aad4fa7 · outbound

This paper cites Negative- aware attention framework for image-text matching,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Negative- aware attention framework for image-text matching,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.464004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.242506Z digest=sha256:e0affa5d3e1cc5ef5e7bc483773b29969a0ba901bd0203f4db5307b14c25fd24

Observation 0c59de8c-3129-4e56-ab57-d752aae450ae · outbound

This paper cites Cyclip: Cyclic contrastive language-image pretraining,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Cyclip: Cyclic contrastive language-image pretraining,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.447654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.247977Z digest=sha256:dff7f0266b4c6bf4edfce999e591f341ea68a8d766b2842ee8ca6a2c7d0c65ce

Observation 0b6c0d14-0423-41e7-bff8-8fa542e372f8 · outbound

This paper cites Improving Multi-Modal Learning with Uni-Modal Teachers.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Improving Multi-Modal Learning with Uni-Modal Teachers

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T15:42:28.253438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:42:28.253438Z digest=sha256:60ad871cc8dcbc46160a77d24c35e14786c715c74f89b6b0a131580508c13c12

Observation 754cb223-4ff8-49d4-a972-b414c9b0019d · outbound

This paper cites Neighborhood discriminant hashing for large-scale image retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Neighborhood discriminant hashing for large-scale image retrieval,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.430514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.258989Z digest=sha256:741256fb5ad8f6c669b1777fae19cf2fb36b571ee39d14556c938a9cee335e1e

Observation f701a0fe-790c-402e-8af7-22e42d96a677 · outbound

This paper cites Learning dis- criminative cross-modality features for rgb-d saliency detection,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Learning dis- criminative cross-modality features for rgb-d saliency detection,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.412090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.264147Z digest=sha256:730bf8ddeabff6c46d7a703f0d9c280a43f0bcad19ea852eaa869197cc27f102

Observation 04bc54b9-f5b4-499b-99a7-cc263718e732 · outbound

This paper cites Camera constraint-free view-based 3-d object retrieval,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Camera constraint-free view-based 3-d object retrieval,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.395834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.268958Z digest=sha256:e4bb82dbe05b71e9b4525c377660da2132ac27117faa5185ba9ecb1c3b78f257

Observation 124c95db-928b-4e4b-a04b-b7521a1177eb · outbound

This paper cites Picture it in your mind: generating high- level visual representations from textual descriptions,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Picture it in your mind: generating high- level visual representations from textual descriptions,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.379885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.273983Z digest=sha256:8515bde9219743db589be5499f350ce6be771320f93bfdde08c4a8a883d86bf8

Observation 104a5a55-ba63-401e-8b87-808f7872fa42 · outbound

This paper cites Attention on attention for image captioning,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Attention on attention for image captioning,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.362920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.279004Z digest=sha256:ec8e367895d027636d56e97929ccd3b63e41a300a60d63423db13ff97ede4f66

Observation bdac5ead-e5a5-4b5c-bab6-46a6d67db623 · outbound

This paper cites Decoupled weight decay regularization,.

Rebalanced Vision-Language Retrieval Considering Structure-Aware Distillation Decoupled weight decay regularization,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:42:28.345018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T15:42:28.284177Z digest=sha256:abb356fb02223c4cff0230b6c6c473963d314e3aaf53ee14d5d0d74c6a66a780

Pith citing papers

No inbound Pith citation observations are available.