Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:28:39.810901Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 100 of 100 outbound references and 0 inbound Pith citation observations for arXiv:2412.01370.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:28:39.810901Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 100 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7ab34dd0-f126-446a-9a2e-6666b0b96ed3 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Artemis: Affective language for visual art
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a559219f-0e2e-4a95-92b1-c3bc742964ff · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Feelingblue: A corpus for understanding the emotional con- notation of color in context.Transactions of the Association for Computational Linguistics, 11:176–190, 2023
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a32c6281-9c4d-4ce6-9854-f0f7ed3dc9aa · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Vqa: Visual question answering
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74832a63-777f-4fca-b59a-95b253eaf969 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Explain me the painting: Multi-topic knowledgeable art description gen- eration
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19fa2347-3eef-48e7-a2c2-1f85b6db6be8 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b79c97ce-f162-480e-8e25-1f1a355f36a8 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Bridg- ing the gap between object and image-level representations for open-vocabulary detection.Advances in Neural Informa- tion Processing Systems, 35:33781–33794, 2022
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2a3c8a7-0c14-412d-8e15-50deef6273e0 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Clip retrieval: Easily compute clip embeddings and build a clip retrieval system with them.https : / / github
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 217bea4a-3900-4979-b31d-fe583452871a · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Viscounth: A large-scale multilin- gual visual question answering dataset for cultural heritage
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 968a67bb-7162-4c82-855b-8b34d7154f9c · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Predicting image aesthetics with deep learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 686e0b9d-99aa-4b79-8619-5920307db9ed · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Vizwiz: nearly real-time answers to visual questions
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e287ec9c-024c-42f7-bbe0-239a78dbecc4 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Visual question answering for cul- tural heritage
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a943caab-bd28-4bf6-b97e-9e6b98f24f67 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Fine-tuning convolutional neural networks for fine art classification.Ex- pert Systems with Applications, 114:107–118, 2018
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f98f85-fceb-4b7b-8d52-d5e9a64d7d93 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Uniter: Universal image-text representation learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd62fb5a-e778-4519-af9e-5bc7a3bf1f7d · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Clip-art: Contrastive pre-training for fine-grained art classification
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 993ed531-a51d-4f56-b838-f2b835c78e60 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Learning sample difficulty from pre-trained models for reliable prediction.Advances in Neural Information Process- ing Systems, 36, 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e79656-ab91-4148-8fd7-ecace7141253 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Novel datasets for fine-grained image categoriza- tion
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc951444-420d-4949-908f-6d3910591cc0 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Noisyart: A dataset for webly-supervised art- work recognition
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e23ee05-b3d7-4684-a88d-373dbc491836 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Webly-supervised zero-shot learning for artwork instance recognition.Pattern Recognition Letters, 128:420– 426, 2019
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c87627a0-b92f-44e6-a460-ca8043706bd8 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Imagenet: A large-scale hierarchical image database
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39835090-6393-42d0-a6f4-58b124e00f7b · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Stytr2: Im- age style transfer with transformers
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5dd0a3b-ad59-40fd-869c-379dad4de782 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning De- coupling zero-shot semantic segmentation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba65d43f-f6d0-4044-a369-8bacfcb9f522 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning A survey on bias in visual datasets.Computer Vision and Image Understanding, 223: 103552, 2022
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d704fb76-e5b4-48f7-95a1-e172a16b8c4b · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Are you talking to a machine? dataset and methods for multilingual image question.Advances in neural information processing systems, 28, 2015
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9dff17b3-abb8-4c4b-9996-d23ab488b45c · outbound
Understanding Museum Exhibits using Vision-Language Reasoning How to read paintings: semantic art understanding with multi-modal retrieval
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 37e29f7e-c594-4e44-b004-7ff0a3c51795 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Knowit vqa: Answering knowledge-based ques- tions about videos
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 93fc2cf7-a3cf-4e9e-9e18-10fcfd7a18e8 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning A dataset and baselines for visual question answering on art
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0aa7eadb-4c48-4b01-8426-e79ec7ed97f1 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning A dataset and baselines for visual question answering on art
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ba39c09d-9e5d-43fb-92dd-f939e800650e · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Im- age style transfer using convolutional neural networks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd7ee42a-0928-46e3-976c-8a75fbd95d3d · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Aes- thetic image captioning from weakly-labelled photographs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 73516bb5-921b-4289-be5d-a6e62b090371 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Beyond language bias: Over- coming multimodal shortcut and distribution biases for ro- bust visual question answering
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b4a5b2ae-22b6-481f-a7b0-5a8031636e2c · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d8dea0-41d4-4287-a82d-315618f7e9ca · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Many- modalqa: Modality disambiguation and qa over diverse in- puts
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1e2af5e0-222f-4f25-bbba-534691a917ab · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Scaling up visual and vision-language representa- tion learning with noisy text supervision
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9ac5b94-5043-4e0c-81d4-d2127b34d9d7 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 014e6e6c-03ae-4007-be21-1b55ee5b5f38 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Prompting visual-language models for efficient video understanding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 56a5d640-b34e-4d6f-9d64-48d150e4c2b6 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning FigureQA: An Annotated Figure Dataset for Visual Reasoning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91a5111b-56c1-408a-b851-afc328280b05 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Mdetr- modulated detection for end-to-end multi-modal understand- ing
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e9ad3cc9-660a-4927-9144-b1305de58ffc · outbound
Understanding Museum Exhibits using Vision-Language Reasoning From word embeddings to document distances
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f243837c-0604-4e8e-8bc2-d85da5c838bb · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Clipstyler: Image style transfer with a single text condition
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a4071822-4b8d-42fe-8d76-44e138dbacfa · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Language-driven semantic seg- mentation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 354d88fa-3e2b-4e3a-b2e8-9ba7caf022bc · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Language-driven Semantic Segmentation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c046909-4675-4dbd-aa31-1c2964475703 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9753162f-58c1-49f5-b48e-d55cf528f2d4 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning VisualBERT: A Simple and Performant Baseline for Vision and Language
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9c55e5e-e360-4186-b272-66c623722a76 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Oscar: Object-semantics aligned pre-training for vision-language tasks
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 668da33f-32df-46d9-9636-97f15eca9a50 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Microsoft coco: Common objects in context
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acf6eea0-eca4-4a60-8565-670eac19241c · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Fine-grained late-interaction multi-modal retrieval for retrieval augmented visual question answering
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 098a63f3-b619-4544-9194-ea597e1a9687 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Visual instruction tuning, 2023
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c81f9129-ed1e-4b7d-a392-f2eea8294171 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Decoupled Weight Decay Regularization
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 003cb9a8-bf1f-47de-a4f8-cd0c19128723 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.Advances in neural information processing systems, 32, 2019
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cbcad21-eaf7-44e9-8c4d-a5ae0d4fd08e · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Data- efficient image captioning of fine art paintings via virtual- real semantic alignment training.Neurocomputing, 490:163– 180, 2022
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6cdcf8fc-e786-4aae-83d4-5b6c3ff86082 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Class-agnostic object detection with multi- modal transformer
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cbfab324-7c10-42f4-b27a-7158d197c08a · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Fine-Grained Visual Classification of Aircraft
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0748b918-e7fc-4108-9cd0-4ce397e63939 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning A multi-world ap- proach to question answering about real-world scenes based on uncertain input.Advances in neural information process- ing systems, 27, 2014
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e6246572-4564-4010-9ec5-c189df868763 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Ask your neurons: A neural-based approach to answering questions about images
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3c9a29dc-8cc1-4c7a-81d1-7531986fefe4 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eb64fad0-9d73-4b02-86f3-debbafd07243 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Taylor & Francis, 2008
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a19689f0-cd54-4a7e-abc3-667faf400d9a · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Foundation Model is Efficient Multimodal Multitask Model Selector
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f9aafe1e-cd7c-4ad7-b6d2-8e9df7233284 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning The rijksmuseum challenge: Museum-centered visual recognition
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b784edc1-0dd0-4ff1-a425-551c1f242782 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 26c6844f-aa70-497a-9df4-85d53dd1cb30 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning A dataset and a con- volutional model for iconography classification in paintings
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 51c3790b-0b65-4663-b441-bc3ac2f5e2e5 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Expanding language-image pretrained models for gen- eral video recognition
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8b3dc484-7cf8-4189-9d71-02bb7a0dd433 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning A survey of geospatial semantic web for cultural heritage
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7482a97c-7bc9-4211-b6ac-3a571966d964 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Suppressing biased samples for robust vqa.IEEE Transactions on Multimedia, 24:3405– 3415, 2021
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dc5d31ef-e9c6-4d4b-9511-18ae44cab577 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Bleu: a method for automatic evaluation of machine translation
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b795c7f2-ee04-4d8e-a045-5abfd8ee07a3 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Combined scal- ing for zero-shot transfer learning.Neurocomputing, 555: 126658, 2023
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b3ce181e-f90d-4a03-993e-6941ebef5c9b · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Learning transferable visual models from natural language supervi- sion
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3393842a-f175-4673-bf18-b1edf5d1071e · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Zero-shot text-to-image generation
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74699bd3-f761-4f52-9766-8131ff3b5a6e · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Fine-tuned clip models are efficient video learners
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dee15f97-4987-4f51-a781-fd26ff89b7d8 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Stylebabel: Artistic style tag- ging and captioning
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c979b1ca-624e-4080-8416-cca8ccfa4758 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Viske: Visual knowledge extraction and question answering by visual verification of relation phrases
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c7124083-46a9-4827-bd99-77b995396922 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning A dataset for multimodal question answering in the cultural heritage domain
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e259545d-07a7-4e8d-b1c1-de2dc9a6d980 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Towards vqa models that can read
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8483041-520a-4f3e-835b-c6508cc17e72 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Bioclip: A vision foundation model for the tree of life
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bba0832d-4e52-4df3-b4e9-8017004c2042 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Omniart: a large- scale artistic benchmark.ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 14 (4):1–21, 2018
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a9614f6d-52ae-4f24-bfa0-b516f6e0cca1 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning MultiModalQA: Complex Question Answering over Text, Tables and Images
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c06e7cdc-e446-442e-b6e8-a092dee5fe18 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Ceci n’est pas une pipe: A deep convo- lutional network for fine-art paintings classification
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cd9533cc-20c9-4381-b957-369923c57141 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Gemini: A Family of Highly Capable Multimodal Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f73830ac-b551-4adf-b5fa-53ed64e0ebd8 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning The caltech-ucsd birds-200–2011 dataset
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9dbba161-2b99-46c5-a7e2-96ae1cca3792 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning ActionCLIP: A New Paradigm for Video Action Recognition
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3abc7c45-00ef-4af2-a55a-71ade113b9d1 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Explicit Knowledge-based Reasoning for Visual Question Answering
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5deeac1-c870-458c-aa33-2d7f6d327429 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Fvqa: Fact-based visual question an- swering.IEEE transactions on pattern analysis and machine intelligence, 40(10):2413–2427, 2017
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d8ad7152-db21-408c-8a55-867df9072f73 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning MedCLIP: Contrastive Learning from Unpaired Medical Images and Text
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bde93852-87da-444a-9af4-350c3fbebaab · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Im- proving clip fine-tuning performance
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3e1bf36c-3379-4ffa-80aa-57ff33191ce3 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Bam! the behance artistic media dataset for recognition beyond photography
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ad4bd166-652f-48e3-a8b3-01743e97da80 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Ask me anything: Free-form vi- sual question answering based on knowledge from external sources
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c61d608d-9033-4403-bef2-4b8ff3d2fab9 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Language bias in Visual Question Answering: A Survey and Taxonomy
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9481fe34-7874-4896-803e-f8985656b4b9 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Lit: Zero-shot transfer with locked-image text tuning
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6333fb85-8626-4a9a-b17e-2f6f341026e1 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning The iMet Collection 2019 Challenge Dataset
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6f322d1f-f461-4a98-a33e-59c7ec488409 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0795a025-0709-4b0d-a524-20aa8616c062 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Extract free dense labels from clip
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f0398813-e35d-4c93-a692-fdfb94ee4b49 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Conditional prompt learning for vision-language mod- els
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80673caf-d90d-4d30-a895-4da7219100f5 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Learning to prompt for vision-language models.In- ternational Journal of Computer Vision, 130(9):2337–2348,
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7af21584-2ef2-4c9b-88e2-43a4b8a837ac · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Detecting twenty-thousand classes using image-level supervision
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 76d54b7f-72ef-45a1-a906-56d441fc0fe2 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Visual7w: Grounded question answering in images
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8433939e-def4-4455-b368-f57453b87aea · outbound
Understanding Museum Exhibits using Vision-Language Reasoning LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e1d174e-af6e-49ce-ba92-ccce016abe46 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning • These aggregators provide access to extensive digitized collections from major museums across Europe and America and offer structured data through platform- specific APIs
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ebe8df6b-8dac-4910-a943-ebc0d528d8f1 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Curation in- volved minimal edits: removing redundant attributes (in- ventory numbers, bibliographic info); extraneous symbols and numbers
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 91b95f06-7661-4bdb-bbf8-2a5fae826243 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Unresolved cited work
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 68c51512-01ec-4ae6-99e5-3cc2bb9dce3b · outbound
Understanding Museum Exhibits using Vision-Language Reasoning Which primary material is the object made of?
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fc021a6f-b5f5-4beb-af3d-2e790aefa241 · outbound
Understanding Museum Exhibits using Vision-Language Reasoning For each object, we now have a list of images and a set of question-answer pairs, omitting the answers for which the value is not known
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
No inbound Pith citation observations are available.