Pith. sign in

Paper Citation Record · LEDGER

Understanding Museum Exhibits using Vision-Language Reasoning

As of 19 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 0 inbound Pith citation observations for arXiv:2412.01370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01370 v2

Coverage vector

measured 100 of 100 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:28:39.810901Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 100 outbound references displayed

  • verified exact4
  • verified fuzzy49
  • unresolved46
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7ab34dd0-f126-446a-9a2e-6666b0b96ed3 · outbound

This paper cites Artemis: Affective language for visual art.

Understanding Museum Exhibits using Vision-Language Reasoning Artemis: Affective language for visual art

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.003322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.003322Z digest=sha256:536fb4b6dd1d9142bc37c52220bbc917de764d9112423d0c8aad91faced32602

Observation a559219f-0e2e-4a95-92b1-c3bc742964ff · outbound

This paper cites Feelingblue: A corpus for understanding the emotional con- notation of color in context.Transactions of the Association for Computational Linguistics, 11:176–190, 2023.

Understanding Museum Exhibits using Vision-Language Reasoning Feelingblue: A corpus for understanding the emotional con- notation of color in context.Transactions of the Association for Computational Linguistics, 11:176–190, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.054864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.054864Z digest=sha256:ead5495598e02956925fc8e7e096dc439f81a54f056b9f664eeb64c538a6c437

Observation a32c6281-9c4d-4ce6-9854-f0f7ed3dc9aa · outbound

This paper cites Vqa: Visual question answering.

Understanding Museum Exhibits using Vision-Language Reasoning Vqa: Visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.145741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.145741Z digest=sha256:e50bfd65c5cc89e1dc7923b3ca1ca355c34ce96b8f2b9426626e9ae51813cee5

Observation 74832a63-777f-4fca-b59a-95b253eaf969 · outbound

This paper cites Explain me the painting: Multi-topic knowledgeable art description gen- eration.

Understanding Museum Exhibits using Vision-Language Reasoning Explain me the painting: Multi-topic knowledgeable art description gen- eration

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.150738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.150738Z digest=sha256:f5be80a4d0ed01fb4f85f92bfb3ff13a521bcbfe6261f078e17fae5934ff0539

Observation 19fa2347-3eef-48e7-a2c2-1f85b6db6be8 · outbound

This paper cites Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits.

Understanding Museum Exhibits using Vision-Language Reasoning Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:28:40.472696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.155517Z digest=sha256:34880b25815ccc464efe57c742e72ef5caa2c44e05fd55f0cc15df371089823a

Observation b79c97ce-f162-480e-8e25-1f1a355f36a8 · outbound

This paper cites Bridg- ing the gap between object and image-level representations for open-vocabulary detection.Advances in Neural Informa- tion Processing Systems, 35:33781–33794, 2022.

Understanding Museum Exhibits using Vision-Language Reasoning Bridg- ing the gap between object and image-level representations for open-vocabulary detection.Advances in Neural Informa- tion Processing Systems, 35:33781–33794, 2022

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.161415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.161415Z digest=sha256:2be06ae216d447596bb3c08f0495f8cdd6396a72c8338c868d2a37c197b1f8e4

Observation d2a3c8a7-0c14-412d-8e15-50deef6273e0 · outbound

This paper cites Clip retrieval: Easily compute clip embeddings and build a clip retrieval system with them.https : / / github.

Understanding Museum Exhibits using Vision-Language Reasoning Clip retrieval: Easily compute clip embeddings and build a clip retrieval system with them.https : / / github

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.166687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.166687Z digest=sha256:422bd3cecac622ba9163cc8417093c8cf7cdd6b81b372c33896bb58c8dce7c45

Observation 217bea4a-3900-4979-b31d-fe583452871a · outbound

This paper cites Viscounth: A large-scale multilin- gual visual question answering dataset for cultural heritage.

Understanding Museum Exhibits using Vision-Language Reasoning Viscounth: A large-scale multilin- gual visual question answering dataset for cultural heritage

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.171052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.171052Z digest=sha256:236afdc1c403c7647ffaed57952e8a7f3f015c0d2a63bdeecfd2a68be14364c4

Observation 968a67bb-7162-4c82-855b-8b34d7154f9c · outbound

This paper cites Predicting image aesthetics with deep learning.

Understanding Museum Exhibits using Vision-Language Reasoning Predicting image aesthetics with deep learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.176728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.176728Z digest=sha256:02d07a3c39bea25a19fa0d8421f422e1e5d68e4496ff6e03d2fc809b816a2b64

Observation 686e0b9d-99aa-4b79-8619-5920307db9ed · outbound

This paper cites Vizwiz: nearly real-time answers to visual questions.

Understanding Museum Exhibits using Vision-Language Reasoning Vizwiz: nearly real-time answers to visual questions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.181314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.181314Z digest=sha256:3544c4c5ec340e615608692f1d06c96e1f0d9e74993c2c65b40dd68db0967a2a

Observation e287ec9c-024c-42f7-bbe0-239a78dbecc4 · outbound

This paper cites Visual question answering for cul- tural heritage.

Understanding Museum Exhibits using Vision-Language Reasoning Visual question answering for cul- tural heritage

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.233064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.233064Z digest=sha256:b8c19daa638608d9861ed166792283873c2a09cd5d4764d2b4811bec495a13f2

Observation a943caab-bd28-4bf6-b97e-9e6b98f24f67 · outbound

This paper cites Fine-tuning convolutional neural networks for fine art classification.Ex- pert Systems with Applications, 114:107–118, 2018.

Understanding Museum Exhibits using Vision-Language Reasoning Fine-tuning convolutional neural networks for fine art classification.Ex- pert Systems with Applications, 114:107–118, 2018

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.318618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.318618Z digest=sha256:4fd803ef0020214e839c5b222427b96bd2d04955c0128ee7484598bb0b74520c

Observation 90f98f85-fceb-4b7b-8d52-d5e9a64d7d93 · outbound

This paper cites Uniter: Universal image-text representation learning.

Understanding Museum Exhibits using Vision-Language Reasoning Uniter: Universal image-text representation learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.402518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.402518Z digest=sha256:61851f8f7d89c75531cfc6553ca46fc24f5a5879291539e5a6a5470f7263637e

Observation dd62fb5a-e778-4519-af9e-5bc7a3bf1f7d · outbound

This paper cites Clip-art: Contrastive pre-training for fine-grained art classification.

Understanding Museum Exhibits using Vision-Language Reasoning Clip-art: Contrastive pre-training for fine-grained art classification

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.423884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.423884Z digest=sha256:cbd208737624ccb048f0712ca36eec48c57b9998ba8345961b6563c346dcfc9a

Observation 993ed531-a51d-4f56-b838-f2b835c78e60 · outbound

This paper cites Learning sample difficulty from pre-trained models for reliable prediction.Advances in Neural Information Process- ing Systems, 36, 2024.

Understanding Museum Exhibits using Vision-Language Reasoning Learning sample difficulty from pre-trained models for reliable prediction.Advances in Neural Information Process- ing Systems, 36, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.428640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.428640Z digest=sha256:f68e84d8208f91ae899ef11ac76d6178b435b3e5e2826e8730cab2e3acbf7997

Observation 45e79656-ab91-4148-8fd7-ecace7141253 · outbound

This paper cites Novel datasets for fine-grained image categoriza- tion.

Understanding Museum Exhibits using Vision-Language Reasoning Novel datasets for fine-grained image categoriza- tion

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.433206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.433206Z digest=sha256:ca88e726a32a653e9b0717acb4567c16e03e94ad53209a096cf2ab2a4a226d4b

Observation dc951444-420d-4949-908f-6d3910591cc0 · outbound

This paper cites Noisyart: A dataset for webly-supervised art- work recognition.

Understanding Museum Exhibits using Vision-Language Reasoning Noisyart: A dataset for webly-supervised art- work recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.438200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.438200Z digest=sha256:d04f61039aae4d6545316a699bf6f0abdec80a8496e7b6abe6ee5eafbd09697a

Observation 3e23ee05-b3d7-4684-a88d-373dbc491836 · outbound

This paper cites Webly-supervised zero-shot learning for artwork instance recognition.Pattern Recognition Letters, 128:420– 426, 2019.

Understanding Museum Exhibits using Vision-Language Reasoning Webly-supervised zero-shot learning for artwork instance recognition.Pattern Recognition Letters, 128:420– 426, 2019

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.443109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.443109Z digest=sha256:464959cc8a8f4bdb2e55afd73f4e9d8deecf8faa5b72ee92d543b85bf05a417e

Observation c87627a0-b92f-44e6-a460-ca8043706bd8 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Understanding Museum Exhibits using Vision-Language Reasoning Imagenet: A large-scale hierarchical image database

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.448126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.448126Z digest=sha256:043518955af0bd8cf4d170dd66541b39cccd8874619809bbf2360ed2923ebaa0

Observation 39835090-6393-42d0-a6f4-58b124e00f7b · outbound

This paper cites Stytr2: Im- age style transfer with transformers.

Understanding Museum Exhibits using Vision-Language Reasoning Stytr2: Im- age style transfer with transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.452791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.452791Z digest=sha256:8212707d9e7dd76cf841131c99fd24ff1a11cb15a1fc33e9300814048c93e47d

Observation a5dd0a3b-ad59-40fd-869c-379dad4de782 · outbound

This paper cites De- coupling zero-shot semantic segmentation.

Understanding Museum Exhibits using Vision-Language Reasoning De- coupling zero-shot semantic segmentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.458317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.458317Z digest=sha256:cca11207a8c9fa37b35830bf28589c790026cd0eb5b9adb7176002b18392a9f7

Observation ba65d43f-f6d0-4044-a369-8bacfcb9f522 · outbound

This paper cites A survey on bias in visual datasets.Computer Vision and Image Understanding, 223: 103552, 2022.

Understanding Museum Exhibits using Vision-Language Reasoning A survey on bias in visual datasets.Computer Vision and Image Understanding, 223: 103552, 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.808497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.462791Z digest=sha256:5059866eb021a5dd8ae6207798ee7a4ddce585f545e202da2488848d2d6b5b4e

Observation d704fb76-e5b4-48f7-95a1-e172a16b8c4b · outbound

This paper cites Are you talking to a machine? dataset and methods for multilingual image question.Advances in neural information processing systems, 28, 2015.

Understanding Museum Exhibits using Vision-Language Reasoning Are you talking to a machine? dataset and methods for multilingual image question.Advances in neural information processing systems, 28, 2015

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.625655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.525414Z digest=sha256:c5d93f22989582a7ff6fab83b4b867de008543f75f2c6664378c05a39efbdabf

Observation 9dff17b3-abb8-4c4b-9996-d23ab488b45c · outbound

This paper cites How to read paintings: semantic art understanding with multi-modal retrieval.

Understanding Museum Exhibits using Vision-Language Reasoning How to read paintings: semantic art understanding with multi-modal retrieval

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.611046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.585444Z digest=sha256:2be3980e960fce7912371d3f230f701e0878a1e2249f39cc244f3b425353665a

Observation 37e29f7e-c594-4e44-b004-7ff0a3c51795 · outbound

This paper cites Knowit vqa: Answering knowledge-based ques- tions about videos.

Understanding Museum Exhibits using Vision-Language Reasoning Knowit vqa: Answering knowledge-based ques- tions about videos

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.594896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.642759Z digest=sha256:dc2bbdf83aa417f0a9041aaefddf9be569fcf8c368c18519e186eeb385dc84b2

Observation 93fc2cf7-a3cf-4e9e-9e18-10fcfd7a18e8 · outbound

This paper cites A dataset and baselines for visual question answering on art.

Understanding Museum Exhibits using Vision-Language Reasoning A dataset and baselines for visual question answering on art

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.547921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.647582Z digest=sha256:e9da5811352755b34cac36abae7cb7cd10b26e08398944c3b28437dc8597ea22

Observation 0aa7eadb-4c48-4b01-8426-e79ec7ed97f1 · outbound

This paper cites A dataset and baselines for visual question answering on art.

Understanding Museum Exhibits using Vision-Language Reasoning A dataset and baselines for visual question answering on art

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.421716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.652677Z digest=sha256:614c0908b1072be7b834900b20069b0f81cd6f96060d921cf23d579c835f5197

Observation ba39c09d-9e5d-43fb-92dd-f939e800650e · outbound

This paper cites Im- age style transfer using convolutional neural networks.

Understanding Museum Exhibits using Vision-Language Reasoning Im- age style transfer using convolutional neural networks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.656505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.656505Z digest=sha256:814c9c46f80b09c713b2a3398e5aaffe4ae0d6cfe03ea554c9afcc0de8b53575

Observation bd7ee42a-0928-46e3-976c-8a75fbd95d3d · outbound

This paper cites Aes- thetic image captioning from weakly-labelled photographs.

Understanding Museum Exhibits using Vision-Language Reasoning Aes- thetic image captioning from weakly-labelled photographs

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.394999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.660234Z digest=sha256:65888086f686f9321e4dc0b581b26bae50bfc95806f1157bd3b52f701f3543cc

Observation 73516bb5-921b-4289-be5d-a6e62b090371 · outbound

This paper cites Beyond language bias: Over- coming multimodal shortcut and distribution biases for ro- bust visual question answering.

Understanding Museum Exhibits using Vision-Language Reasoning Beyond language bias: Over- coming multimodal shortcut and distribution biases for ro- bust visual question answering

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.378533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.664520Z digest=sha256:2a9777cd91f1994a79a36bf6c4efe0897e99811b82a6b559a04f6deb2e6958b6

Observation b4a5b2ae-22b6-481f-a7b0-5a8031636e2c · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Understanding Museum Exhibits using Vision-Language Reasoning Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.668359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.668359Z digest=sha256:0026d1523c169dc870ae39d8b365ab328ee87ac79bc48547b9a1ccd3b5b2483f

Observation f4d8dea0-41d4-4287-a82d-315618f7e9ca · outbound

This paper cites Many- modalqa: Modality disambiguation and qa over diverse in- puts.

Understanding Museum Exhibits using Vision-Language Reasoning Many- modalqa: Modality disambiguation and qa over diverse in- puts

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.247869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.673304Z digest=sha256:bbdb2d8a9aa9a8ad7e3d75a12d0cee64379ad029d031d1f815690494cb4cb545

Observation 1e2af5e0-222f-4f25-bbba-534691a917ab · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Understanding Museum Exhibits using Vision-Language Reasoning Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.677467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.677467Z digest=sha256:a82a48bcc294e4d91403a68c718c9f64a90433037b93ac2e29ef690872aafd1a

Observation f9ac5b94-5043-4e0c-81d4-d2127b34d9d7 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Understanding Museum Exhibits using Vision-Language Reasoning Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.681754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.681754Z digest=sha256:10fe8d4d702c1c751bcca64dc286083d2a3814f32cdf468785b6af92371090e5

Observation 014e6e6c-03ae-4007-be21-1b55ee5b5f38 · outbound

This paper cites Prompting visual-language models for efficient video understanding.

Understanding Museum Exhibits using Vision-Language Reasoning Prompting visual-language models for efficient video understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.186286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.689495Z digest=sha256:e03ae456e8588d95b7ee0f242a6cd065ce74bd378536158cc14917efaeb7b753

Observation 56a5d640-b34e-4d6f-9d64-48d150e4c2b6 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

Understanding Museum Exhibits using Vision-Language Reasoning FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.693551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.693551Z digest=sha256:eabdbc7d50f02176c5f1c521010116ab0262bda99b256ee69a9b8c7f5b0a8df5

Observation 91a5111b-56c1-408a-b851-afc328280b05 · outbound

This paper cites Mdetr- modulated detection for end-to-end multi-modal understand- ing.

Understanding Museum Exhibits using Vision-Language Reasoning Mdetr- modulated detection for end-to-end multi-modal understand- ing

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.169881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.751828Z digest=sha256:0c07119aea217f2da9620d583d130aeda69e5921c9b147e41f41b6f2674b87e9

Observation e9ad3cc9-660a-4927-9144-b1305de58ffc · outbound

This paper cites From word embeddings to document distances.

Understanding Museum Exhibits using Vision-Language Reasoning From word embeddings to document distances

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:43.066016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.844763Z digest=sha256:844df03dc789feae275c9da0d6bc375c39460e45f45cde32a55fb35b9acad687

Observation f243837c-0604-4e8e-8bc2-d85da5c838bb · outbound

This paper cites Clipstyler: Image style transfer with a single text condition.

Understanding Museum Exhibits using Vision-Language Reasoning Clipstyler: Image style transfer with a single text condition

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.938721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.883392Z digest=sha256:d102dd999dff8ec5e340d0f397346c2498f8a25fb85b4175cfb1e0d73d9e5b55

Observation a4071822-4b8d-42fe-8d76-44e138dbacfa · outbound

This paper cites Language-driven semantic seg- mentation.

Understanding Museum Exhibits using Vision-Language Reasoning Language-driven semantic seg- mentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.899404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.888272Z digest=sha256:3936dafeb5a3972a8ee68e0f044c3a002aaca76fd3d9d59a6658833324a29307

Observation 354d88fa-3e2b-4e3a-b2e8-9ba7caf022bc · outbound

This paper cites Language-driven Semantic Segmentation.

Understanding Museum Exhibits using Vision-Language Reasoning Language-driven Semantic Segmentation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.892694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.892694Z digest=sha256:3d870511352124c1bc7a8305cb7f8241fdcf01ccd5101d128bf7dcba761039c9

Observation 6c046909-4675-4dbd-aa31-1c2964475703 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Understanding Museum Exhibits using Vision-Language Reasoning Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.883668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.897377Z digest=sha256:777549634121e9b1f842cbea3382663b05e1f604cdc1a059a5d80f1e00b6a366

Observation 9753162f-58c1-49f5-b48e-d55cf528f2d4 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Understanding Museum Exhibits using Vision-Language Reasoning VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.901539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.901539Z digest=sha256:c281fa343745f4f3ad610d49cd102b8e69bcc58e04c5b205ced719047fb91388

Observation e9c55e5e-e360-4186-b272-66c623722a76 · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

Understanding Museum Exhibits using Vision-Language Reasoning Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.906247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.906247Z digest=sha256:bb143b1deb3ff1adf11e7285f6510c65f7873e405e68c21c564953117bbd8eba

Observation 668da33f-32df-46d9-9636-97f15eca9a50 · outbound

This paper cites Microsoft coco: Common objects in context.

Understanding Museum Exhibits using Vision-Language Reasoning Microsoft coco: Common objects in context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.910224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.910224Z digest=sha256:5d325bee737cb8df5bd793923a4f080beac084712bfb601ea033e6bb2944fa97

Observation acf6eea0-eca4-4a60-8565-670eac19241c · outbound

This paper cites Fine-grained late-interaction multi-modal retrieval for retrieval augmented visual question answering.

Understanding Museum Exhibits using Vision-Language Reasoning Fine-grained late-interaction multi-modal retrieval for retrieval augmented visual question answering

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.844970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.914415Z digest=sha256:14d9a2412d98787da0bb861a1a081f3a733dc0650c23555857fb8bda0eea5c94

Observation 098a63f3-b619-4544-9194-ea597e1a9687 · outbound

This paper cites Visual instruction tuning, 2023.

Understanding Museum Exhibits using Vision-Language Reasoning Visual instruction tuning, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.717687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:38.918582Z digest=sha256:a70dd3d17b34ec3492216ad0a455120c2a4b0cc2d6a8d2cdf054b2c70d8f59ab

Observation c81f9129-ed1e-4b7d-a392-f2eea8294171 · outbound

This paper cites Decoupled Weight Decay Regularization.

Understanding Museum Exhibits using Vision-Language Reasoning Decoupled Weight Decay Regularization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:38.956848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:38.956848Z digest=sha256:0356742e526fd26f5713bdf74947f03241d3c4838f5e1a366ebaa74bf84eec4a

Observation 003cb9a8-bf1f-47de-a4f8-cd0c19128723 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.Advances in neural information processing systems, 32, 2019.

Understanding Museum Exhibits using Vision-Language Reasoning Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.Advances in neural information processing systems, 32, 2019

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.001582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.001582Z digest=sha256:dc5ee44e91ede16a5ffdc819feee4e603c730a730fd488ae8fe25c93e6e0fa32

Observation 8cbcad21-eaf7-44e9-8c4d-a5ae0d4fd08e · outbound

This paper cites Data- efficient image captioning of fine art paintings via virtual- real semantic alignment training.Neurocomputing, 490:163– 180, 2022.

Understanding Museum Exhibits using Vision-Language Reasoning Data- efficient image captioning of fine art paintings via virtual- real semantic alignment training.Neurocomputing, 490:163– 180, 2022

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.635288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.068238Z digest=sha256:ace7f3066d59aa1bb317eb6f5e8c018a3f3640f05fd4e03c8b7cab90bfc9a555

Observation 6cdcf8fc-e786-4aae-83d4-5b6c3ff86082 · outbound

This paper cites Class-agnostic object detection with multi- modal transformer.

Understanding Museum Exhibits using Vision-Language Reasoning Class-agnostic object detection with multi- modal transformer

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.618986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.090766Z digest=sha256:7e8e523b87801f580eeb0fd2f80485f470e39609e2963ecdbd2fc6128c596292

Observation cbfab324-7c10-42f4-b27a-7158d197c08a · outbound

This paper cites Fine-Grained Visual Classification of Aircraft.

Understanding Museum Exhibits using Vision-Language Reasoning Fine-Grained Visual Classification of Aircraft

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.096188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.096188Z digest=sha256:9c43d9f55f687d792c6324b51ae3c9d8821f069a0ed145ba8a7fe9f20f71d73d

Observation 0748b918-e7fc-4108-9cd0-4ce397e63939 · outbound

This paper cites A multi-world ap- proach to question answering about real-world scenes based on uncertain input.Advances in neural information process- ing systems, 27, 2014.

Understanding Museum Exhibits using Vision-Language Reasoning A multi-world ap- proach to question answering about real-world scenes based on uncertain input.Advances in neural information process- ing systems, 27, 2014

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.603957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.100756Z digest=sha256:fec9ba741ac49dadf8d665a2e9145d73befb190f7176c7b314c3edeea3c7a461

Observation e6246572-4564-4010-9ec5-c189df868763 · outbound

This paper cites Ask your neurons: A neural-based approach to answering questions about images.

Understanding Museum Exhibits using Vision-Language Reasoning Ask your neurons: A neural-based approach to answering questions about images

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.588756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.105087Z digest=sha256:1a0b4bee112d9a5c1e73b5e0f4effed98bad37b8b59262dae7da333e2ab96738

Observation 3c9a29dc-8cc1-4c7a-81d1-7531986fefe4 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Understanding Museum Exhibits using Vision-Language Reasoning Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.573678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.109382Z digest=sha256:4d57e2a4e2c052c65c21eb80379d2938d76f705cfca6d0de3d2ac856c6af120e

Observation eb64fad0-9d73-4b02-86f3-debbafd07243 · outbound

This paper cites Taylor & Francis, 2008.

Understanding Museum Exhibits using Vision-Language Reasoning Taylor & Francis, 2008

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.425263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.113376Z digest=sha256:15740b940f4d8a5ab8e65347d9f99838db50c09d6c33edfa91794a1fc698166b

Observation a19689f0-cd54-4a7e-abc3-667faf400d9a · outbound

This paper cites Foundation Model is Efficient Multimodal Multitask Model Selector.

Understanding Museum Exhibits using Vision-Language Reasoning Foundation Model is Efficient Multimodal Multitask Model Selector

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:28:40.216299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.117862Z digest=sha256:d2996804d7a3e2b026a66a8465d85302ac8d8fd77f79265ca1ab18490ccac3b5

Observation f9aafe1e-cd7c-4ad7-b6d2-8e9df7233284 · outbound

This paper cites The rijksmuseum challenge: Museum-centered visual recognition.

Understanding Museum Exhibits using Vision-Language Reasoning The rijksmuseum challenge: Museum-centered visual recognition

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.318571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.123393Z digest=sha256:3df37a95252f15bc721231f325910a306c85efe59f094b22ac7dff50c6cf8717

Observation b784edc1-0dd0-4ff1-a425-551c1f242782 · outbound

This paper cites Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories.

Understanding Museum Exhibits using Vision-Language Reasoning Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.302337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.127890Z digest=sha256:ad18096bd4e2c9877e1740d2c21d35a93d2630a3a7bd7961cb2663bb26a1521a

Observation 26c6844f-aa70-497a-9df4-85d53dd1cb30 · outbound

This paper cites A dataset and a con- volutional model for iconography classification in paintings.

Understanding Museum Exhibits using Vision-Language Reasoning A dataset and a con- volutional model for iconography classification in paintings

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.285300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.132310Z digest=sha256:4eb1f4a5d12a4d9207296c2c98add653ce6443a0bbaadbbb9523e0e7dcf69bb3

Observation 51c3790b-0b65-4663-b441-bc3ac2f5e2e5 · outbound

This paper cites Expanding language-image pretrained models for gen- eral video recognition.

Understanding Museum Exhibits using Vision-Language Reasoning Expanding language-image pretrained models for gen- eral video recognition

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.267591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.167324Z digest=sha256:81a899a42d397c77995e315e880d7850454cf7e46f4da7a3ae61802b0cd33ec4

Observation 8b3dc484-7cf8-4189-9d71-02bb7a0dd433 · outbound

This paper cites A survey of geospatial semantic web for cultural heritage.

Understanding Museum Exhibits using Vision-Language Reasoning A survey of geospatial semantic web for cultural heritage

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:42.072177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.222939Z digest=sha256:236ee24b291a5f2355b7a87d3b088f829936565c37a628bcb839a6e4294df47c

Observation 7482a97c-7bc9-4211-b6ac-3a571966d964 · outbound

This paper cites Suppressing biased samples for robust vqa.IEEE Transactions on Multimedia, 24:3405– 3415, 2021.

Understanding Museum Exhibits using Vision-Language Reasoning Suppressing biased samples for robust vqa.IEEE Transactions on Multimedia, 24:3405– 3415, 2021

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.957607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.285954Z digest=sha256:f9d048a02d7c8f3c19844d9969eb3ab64cd92e5a7e01886e38edbea320756687

Observation dc5d31ef-e9c6-4d4b-9511-18ae44cab577 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Understanding Museum Exhibits using Vision-Language Reasoning Bleu: a method for automatic evaluation of machine translation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.941560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.314365Z digest=sha256:69ca5c74f3038a65534e14917cf7d81a184e5d06af12ec02fb975bb41f030848

Observation b795c7f2-ee04-4d8e-a045-5abfd8ee07a3 · outbound

This paper cites Combined scal- ing for zero-shot transfer learning.Neurocomputing, 555: 126658, 2023.

Understanding Museum Exhibits using Vision-Language Reasoning Combined scal- ing for zero-shot transfer learning.Neurocomputing, 555: 126658, 2023

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.925823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.319230Z digest=sha256:cefa7ece7a866b434a88932aae603f2997b0522f9c8c76777476877b38a2782d

Observation b3ce181e-f90d-4a03-993e-6941ebef5c9b · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Understanding Museum Exhibits using Vision-Language Reasoning Learning transferable visual models from natural language supervi- sion

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.324548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.324548Z digest=sha256:523f8165f7b8f098c3e0b14a46ef1dbdcd8984b3078a9e8181627942bf682331

Observation 3393842a-f175-4673-bf18-b1edf5d1071e · outbound

This paper cites Zero-shot text-to-image generation.

Understanding Museum Exhibits using Vision-Language Reasoning Zero-shot text-to-image generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.330273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.330273Z digest=sha256:a87735ce85be8d0b8afe733a725223c2606c540ad5a5a9a33e7bf254188a4501

Observation 74699bd3-f761-4f52-9766-8131ff3b5a6e · outbound

This paper cites Fine-tuned clip models are efficient video learners.

Understanding Museum Exhibits using Vision-Language Reasoning Fine-tuned clip models are efficient video learners

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.730928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.334548Z digest=sha256:9841f16628c4ea874a9e1a85ce20af5af767247f8f7aaf8074a93dc724006357

Observation dee15f97-4987-4f51-a781-fd26ff89b7d8 · outbound

This paper cites Stylebabel: Artistic style tag- ging and captioning.

Understanding Museum Exhibits using Vision-Language Reasoning Stylebabel: Artistic style tag- ging and captioning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.713807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.339012Z digest=sha256:85e534b16a64c3264a74d72795c770c9e8ae33ae3dca754980e55f30c4b0a3d7

Observation c979b1ca-624e-4080-8416-cca8ccfa4758 · outbound

This paper cites Viske: Visual knowledge extraction and question answering by visual verification of relation phrases.

Understanding Museum Exhibits using Vision-Language Reasoning Viske: Visual knowledge extraction and question answering by visual verification of relation phrases

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.698238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.343060Z digest=sha256:b7a19b419252cca1025c7342b7e3f9eb43e2b3d6257a0beeb08235c0176f4396

Observation c7124083-46a9-4827-bd99-77b995396922 · outbound

This paper cites A dataset for multimodal question answering in the cultural heritage domain.

Understanding Museum Exhibits using Vision-Language Reasoning A dataset for multimodal question answering in the cultural heritage domain

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.559034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.347348Z digest=sha256:21f0549fca31c0b66a2930569f7591908d3cf95c3ce13bc20fa381cad426ffcf

Observation e259545d-07a7-4e8d-b1c1-de2dc9a6d980 · outbound

This paper cites Towards vqa models that can read.

Understanding Museum Exhibits using Vision-Language Reasoning Towards vqa models that can read

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.351442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.351442Z digest=sha256:59d296bd8e5a93dcd8468cade4291298f55225392b2a4f7aafe0a5a468d97c6e

Observation d8483041-520a-4f3e-835b-c6508cc17e72 · outbound

This paper cites Bioclip: A vision foundation model for the tree of life.

Understanding Museum Exhibits using Vision-Language Reasoning Bioclip: A vision foundation model for the tree of life

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.355265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.355265Z digest=sha256:03381eaacc4b637765e88574596a6914c87a335233196206c376438e7bcfc91e

Observation bba0832d-4e52-4df3-b4e9-8017004c2042 · outbound

This paper cites Omniart: a large- scale artistic benchmark.ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 14 (4):1–21, 2018.

Understanding Museum Exhibits using Vision-Language Reasoning Omniart: a large- scale artistic benchmark.ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 14 (4):1–21, 2018

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.384474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.360221Z digest=sha256:8e03efc801df1cfe8a2c2dea5620b20444694a82b296bd3a73a55466fe42168b

Observation a9614f6d-52ae-4f24-bfa0-b516f6e0cca1 · outbound

This paper cites MultiModalQA: Complex Question Answering over Text, Tables and Images.

Understanding Museum Exhibits using Vision-Language Reasoning MultiModalQA: Complex Question Answering over Text, Tables and Images

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.364514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.364514Z digest=sha256:081771ae147b8d78d5cf3f8d6eb863f4143cddd942e49599774346a1387374c1

Observation c06e7cdc-e446-442e-b6e8-a092dee5fe18 · outbound

This paper cites Ceci n’est pas une pipe: A deep convo- lutional network for fine-art paintings classification.

Understanding Museum Exhibits using Vision-Language Reasoning Ceci n’est pas une pipe: A deep convo- lutional network for fine-art paintings classification

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.291712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.369035Z digest=sha256:20ba804b04ad6e69c8527a4f7659d15cea2a82d9f81a5b3beb8e65b715fd681c

Observation cd9533cc-20c9-4381-b957-369923c57141 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Understanding Museum Exhibits using Vision-Language Reasoning Gemini: A Family of Highly Capable Multimodal Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.373214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.373214Z digest=sha256:545eedac7bbaef63c0d076276514c0273610c0472865a169199963413f694c34

Observation f73830ac-b551-4adf-b5fa-53ed64e0ebd8 · outbound

This paper cites The caltech-ucsd birds-200–2011 dataset.

Understanding Museum Exhibits using Vision-Language Reasoning The caltech-ucsd birds-200–2011 dataset

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.178391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.377956Z digest=sha256:265c1a26b6235523c1bc061cd1fd803f101a771a14931475741fcb2a37db8dd7

Observation 9dbba161-2b99-46c5-a7e2-96ae1cca3792 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Understanding Museum Exhibits using Vision-Language Reasoning ActionCLIP: A New Paradigm for Video Action Recognition

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.410964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.410964Z digest=sha256:d7394bdd7816ef030e482764da9be91474e7dc4195bc303437d82f22014bda5a

Observation 3abc7c45-00ef-4af2-a55a-71ade113b9d1 · outbound

This paper cites Explicit Knowledge-based Reasoning for Visual Question Answering.

Understanding Museum Exhibits using Vision-Language Reasoning Explicit Knowledge-based Reasoning for Visual Question Answering

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.481628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.481628Z digest=sha256:2343ba13813aed484c458f157368e756e030c7076774a65988ac37f7a77bdf8e

Observation d5deeac1-c870-458c-aa33-2d7f6d327429 · outbound

This paper cites Fvqa: Fact-based visual question an- swering.IEEE transactions on pattern analysis and machine intelligence, 40(10):2413–2427, 2017.

Understanding Museum Exhibits using Vision-Language Reasoning Fvqa: Fact-based visual question an- swering.IEEE transactions on pattern analysis and machine intelligence, 40(10):2413–2427, 2017

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.162644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.508493Z digest=sha256:f83e121120d237a7c825366409634a1b1ff7be3a5c58980cffa7e63a9f01695d

Observation d8ad7152-db21-408c-8a55-867df9072f73 · outbound

This paper cites MedCLIP: Contrastive Learning from Unpaired Medical Images and Text.

Understanding Museum Exhibits using Vision-Language Reasoning MedCLIP: Contrastive Learning from Unpaired Medical Images and Text

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.513298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.513298Z digest=sha256:f535e48da73f32f593deeb46d2f9aba4994d75297dd8a4d8944eb9e63d370857

Observation bde93852-87da-444a-9af4-350c3fbebaab · outbound

This paper cites Im- proving clip fine-tuning performance.

Understanding Museum Exhibits using Vision-Language Reasoning Im- proving clip fine-tuning performance

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.145829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.519200Z digest=sha256:916830bbab6f908cdf2d3e8237475d7ac802d5d067d235369c4210608b8f0c44

Observation 3e1bf36c-3379-4ffa-80aa-57ff33191ce3 · outbound

This paper cites Bam! the behance artistic media dataset for recognition beyond photography.

Understanding Museum Exhibits using Vision-Language Reasoning Bam! the behance artistic media dataset for recognition beyond photography

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.130513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.524116Z digest=sha256:17b28b093fe459d873c0bdabc96b1f3c39981e7aeae965af0f21bb7c30ea3fee

Observation ad4bd166-652f-48e3-a8b3-01743e97da80 · outbound

This paper cites Ask me anything: Free-form vi- sual question answering based on knowledge from external sources.

Understanding Museum Exhibits using Vision-Language Reasoning Ask me anything: Free-form vi- sual question answering based on knowledge from external sources

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.115755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.528325Z digest=sha256:d61ee6b652af012116d289792c330a0a069deaf86eabdb524d4d4047fa2970af

Observation c61d608d-9033-4403-bef2-4b8ff3d2fab9 · outbound

This paper cites Language bias in Visual Question Answering: A Survey and Taxonomy.

Understanding Museum Exhibits using Vision-Language Reasoning Language bias in Visual Question Answering: A Survey and Taxonomy

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:28:40.036189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.532262Z digest=sha256:7fbff0e3724dada7bebee7af51dd016ea57106aeb76bfd3910d8b12acf9f8608

Observation 9481fe34-7874-4896-803e-f8985656b4b9 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

Understanding Museum Exhibits using Vision-Language Reasoning Lit: Zero-shot transfer with locked-image text tuning

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:41.049754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.537077Z digest=sha256:069a3ba8baf20731b9f822b609263e0a2cc3850df9ff559986d59ade9d71691a

Observation 6333fb85-8626-4a9a-b17e-2f6f341026e1 · outbound

This paper cites The iMet Collection 2019 Challenge Dataset.

Understanding Museum Exhibits using Vision-Language Reasoning The iMet Collection 2019 Challenge Dataset

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:28:39.948829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.541446Z digest=sha256:69a93dce521a22d7b603e91734bf7279f9480d66c1be541d6bbec4d5bcce3c15

Observation 6f322d1f-f461-4a98-a33e-59c7ec488409 · outbound

This paper cites Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling.

Understanding Museum Exhibits using Vision-Language Reasoning Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.545705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.545705Z digest=sha256:6ea9179d424a36e2db2a6bc52094afc25385ae9d81bb24989fb8d64fec7883e6

Observation 0795a025-0709-4b0d-a524-20aa8616c062 · outbound

This paper cites Extract free dense labels from clip.

Understanding Museum Exhibits using Vision-Language Reasoning Extract free dense labels from clip

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:40.902216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.622118Z digest=sha256:989d4dafaca5cb29c57e8836c2c43dfbdbd3db9a35efb2abf2502aeff1ab4377

Observation f0398813-e35d-4c93-a692-fdfb94ee4b49 · outbound

This paper cites Conditional prompt learning for vision-language mod- els.

Understanding Museum Exhibits using Vision-Language Reasoning Conditional prompt learning for vision-language mod- els

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.722009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.722009Z digest=sha256:5c031bf693b8d9f3cad0f6c715e567038f535a60f853c7e43cd10fd9f084ba84

Observation 80673caf-d90d-4d30-a895-4da7219100f5 · outbound

This paper cites Learning to prompt for vision-language models.In- ternational Journal of Computer Vision, 130(9):2337–2348,.

Understanding Museum Exhibits using Vision-Language Reasoning Learning to prompt for vision-language models.In- ternational Journal of Computer Vision, 130(9):2337–2348,

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.774502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.774502Z digest=sha256:df2effcaa06fbc721bc1e14a55d3d68cf82f43cdf7fd3afca1d9f4478558c014

Observation 7af21584-2ef2-4c9b-88e2-43a4b8a837ac · outbound

This paper cites Detecting twenty-thousand classes using image-level supervision.

Understanding Museum Exhibits using Vision-Language Reasoning Detecting twenty-thousand classes using image-level supervision

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:40.867945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.778885Z digest=sha256:0d227a6b1bb09ebf849c1abe7c40f86a5ca807f3e3ecfb4e43847036418d9f2a

Observation 76d54b7f-72ef-45a1-a906-56d441fc0fe2 · outbound

This paper cites Visual7w: Grounded question answering in images.

Understanding Museum Exhibits using Vision-Language Reasoning Visual7w: Grounded question answering in images

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:40.853371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.783992Z digest=sha256:1725fe1c4158b0cfd0dc4237c9880af04498a3eeefd9ec6275e5a33171dcb7aa

Observation 8433939e-def4-4455-b368-f57453b87aea · outbound

This paper cites LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model.

Understanding Museum Exhibits using Vision-Language Reasoning LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T04:28:39.788524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:28:39.788524Z digest=sha256:2d98c03b01e285e719b3bafd87cdab7d04bbaaf16aeab68ae49b6a7a69649037

Observation 9e1d174e-af6e-49ce-ba92-ccce016abe46 · outbound

This paper cites • These aggregators provide access to extensive digitized collections from major museums across Europe and America and offer structured data through platform- specific APIs.

Understanding Museum Exhibits using Vision-Language Reasoning • These aggregators provide access to extensive digitized collections from major museums across Europe and America and offer structured data through platform- specific APIs

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:40.751013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.793770Z digest=sha256:59deb834bbf46a42b46ea4f02840c3edf861bbd0faf10bb93acc65e25a548b46

Observation ebe8df6b-8dac-4910-a943-ebc0d528d8f1 · outbound

This paper cites Curation in- volved minimal edits: removing redundant attributes (in- ventory numbers, bibliographic info); extraneous symbols and numbers.

Understanding Museum Exhibits using Vision-Language Reasoning Curation in- volved minimal edits: removing redundant attributes (in- ventory numbers, bibliographic info); extraneous symbols and numbers

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:40.619143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.798428Z digest=sha256:4053e83a86cbf7ff82a7428babf6c02ac05c51b899be19ce6ede84713dce7685

Observation 91b95f06-7661-4bdb-bbf8-2a5fae826243 · outbound

This paper cites an unresolved cited work.

Understanding Museum Exhibits using Vision-Language Reasoning Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:28:40.604256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.802733Z digest=sha256:364223fdda4d33266cb542ab8e4c28672d2da926ae32aa873d3c471f41d96081

Observation 68c51512-01ec-4ae6-99e5-3cc2bb9dce3b · outbound

This paper cites Which primary material is the object made of?.

Understanding Museum Exhibits using Vision-Language Reasoning Which primary material is the object made of?

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:28:40.587709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.806871Z digest=sha256:da298e25a2c2844aefa8b9f1b152786d13f575a0364ff3164e127975173ae8be

Observation fc021a6f-b5f5-4beb-af3d-2e790aefa241 · outbound

This paper cites For each object, we now have a list of images and a set of question-answer pairs, omitting the answers for which the value is not known.

Understanding Museum Exhibits using Vision-Language Reasoning For each object, we now have a list of images and a set of question-answer pairs, omitting the answers for which the value is not known

Reference 100

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T04:28:40.572404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T04:28:39.810901Z digest=sha256:61f438c50e9bc6f5badc484443509f8ccc7ef2a6b3c5fa867c8b578e2acb4137

Pith citing papers

No inbound Pith citation observations are available.