Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:19:13.746242Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2501.02201.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:19:13.746242Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-11T02:03:52.566413Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T04:00:54.597918Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1db03ffc-a8cd-4cb3-95ea-be41ea326195 · outbound
Acknowledging Focus Ambiguity in Visual Questions Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abb0115e-9ccc-42c3-a432-c218d43d2764 · outbound
Acknowledging Focus Ambiguity in Visual Questions Vqa: Visual question answering
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 046961b5-25f2-4b06-adae-f8c25fdb0829 · outbound
Acknowledging Focus Ambiguity in Visual Questions Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2caccec5-8c0d-4c12-985c-10cf576219e7 · outbound
Acknowledging Focus Ambiguity in Visual Questions Why does a visual question have different answers? In Proceedings of the IEEE International Conference on Computer Vision , pages 4271–4280, 2019
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 35cf7e17-5a7b-4327-bf6b-7dbb48f71430 · outbound
Acknowledging Focus Ambiguity in Visual Questions Grounding answers for visual questions asked by visually impaired people
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5e6765d8-826e-47b3-9dcc-b2f0c189e2f7 · outbound
Acknowledging Focus Ambiguity in Visual Questions Vqa therapy: Exploring answer differences by visually ground- ing answers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 635550d4-1e76-413f-a77b-8615a1694fc8 · outbound
Acknowledging Focus Ambiguity in Visual Questions Fully authentic visual question answering dataset from online communities
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 011efa55-1d4e-4a66-ad1c-9fc2c1d88cf9 · outbound
Acknowledging Focus Ambiguity in Visual Questions Cops-ref: A new dataset and task on composi- tional referring expression comprehension
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a0e3da99-d8f5-4681-a206-039167b74e9e · outbound
Acknowledging Focus Ambiguity in Visual Questions How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9db10d44-fa89-48d2-a68f-fa6e216b624f · outbound
Acknowledging Focus Ambiguity in Visual Questions Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a3bbdbb-24c3-47b5-93be-e98729387ccb · outbound
Acknowledging Focus Ambiguity in Visual Questions Resolving Language and Vision Ambiguities Together: Joint Segmentation & Prepositional Attachment Resolution in Captioned Scenes
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b3bcbc7a-8ea3-4ab5-a549-b2e4cfcff3b4 · outbound
Acknowledging Focus Ambiguity in Visual Questions Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1e47f0a-baf4-4de6-b1bf-343d711dfa2c · outbound
Acknowledging Focus Ambiguity in Visual Questions Zero-shot and few-shot video question answering with multi-modal prompts
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c48e1e52-e66b-416e-9b4d-6f96680ec71c · outbound
Acknowledging Focus Ambiguity in Visual Questions Vqs: Linking segmentations to questions and answers for supervised attention in vqa and question-focused semantic segmentation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 853b4b89-a37d-4f41-8f87-377224909b5a · outbound
Acknowledging Focus Ambiguity in Visual Questions Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 34991dcb-c056-4db1-b3ab-d884437dda1a · outbound
Acknowledging Focus Ambiguity in Visual Questions Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57477597-2063-4653-83f7-ad35ea92d844 · outbound
Acknowledging Focus Ambiguity in Visual Questions Abg-coqa: Clarifying ambiguity in conversa- tional question answering
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 57820aaf-2d32-4f42-b173-60d5980fc62f · outbound
Acknowledging Focus Ambiguity in Visual Questions Lvis: A dataset for large vocabulary instance segmentation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4201e5f5-1fef-4fd7-9c2b-d9caf787616d · outbound
Acknowledging Focus Ambiguity in Visual Questions Crowdverge: Predicting if people will agree on the answer to a visual question
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2e105630-43e1-464d-b41e-f46e7fb7c367 · outbound
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ea230c51-3467-41b8-9204-36ecdbb20ccc · outbound
Acknowledging Focus Ambiguity in Visual Questions Predicting foreground object ambiguity and efficiently crowdsourcing the segmentation (s)
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c4a3b52b-ac85-46c4-8ffb-64246d111d63 · outbound
Acknowledging Focus Ambiguity in Visual Questions Vizwiz grand challenge: Answering visual questions from blind people
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4804eaf0-a390-47bf-a4f2-ea42c7e351c2 · outbound
Acknowledging Focus Ambiguity in Visual Questions A survey on instance segmentation: state of the art.International jour- nal of multimedia information retrieval, 9(3):171–189, 2020
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8a8c6159-872b-4490-b94c-f1783ee6edfa · outbound
Acknowledging Focus Ambiguity in Visual Questions Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a8dfdd65-a7c9-4e4a-b377-45a87af2f0c7 · outbound
Acknowledging Focus Ambiguity in Visual Questions Long-Form Answers to Visual Questions from Blind and Low Vision People
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 850fb64c-d769-4a76-ad0b-9fea91b5cc17 · outbound
Acknowledging Focus Ambiguity in Visual Questions Salient object detection: A discriminative regional feature integration approach
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 46eb0406-f91b-444f-8b4c-68b9d3653621 · outbound
Acknowledging Focus Ambiguity in Visual Questions Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c97303a-e98f-4b4f-9231-252b696bc887 · outbound
Acknowledging Focus Ambiguity in Visual Questions Segment any- thing
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4f380e9f-5f23-40d4-8959-86a1e95516b4 · outbound
Acknowledging Focus Ambiguity in Visual Questions Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 496ad413-1645-422f-8535-747f1edabc8f · outbound
Acknowledging Focus Ambiguity in Visual Questions SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6195d37c-5db3-4e36-a8d1-09fe210799ee · outbound
Acknowledging Focus Ambiguity in Visual Questions Microsoft coco: Common objects in context
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 214bc35a-d889-44c9-9e38-7e83f092d068 · outbound
Acknowledging Focus Ambiguity in Visual Questions Llava-next: Im- proved reasoning, ocr, and world knowledge, January 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0905dfff-9a34-430b-85ea-f7fd155a9d64 · outbound
Acknowledging Focus Ambiguity in Visual Questions Learning to detect a salient object
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 34b8eedb-3687-4a2e-be56-31c07a7ade74 · outbound
Acknowledging Focus Ambiguity in Visual Questions Storytelling with image data: a systematic review and comparative anal- ysis of methods and tools
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation efe7ec67-707a-43b3-a622-7daad9b8fd6c · outbound
Acknowledging Focus Ambiguity in Visual Questions MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62494ba0-6807-4115-b024-579d2c88284d · outbound
Acknowledging Focus Ambiguity in Visual Questions Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 90ee9174-dae7-4426-a98c-5ab9c659b7c8 · outbound
Acknowledging Focus Ambiguity in Visual Questions Resolving ambi- guities in text-to-image generative models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 08c67028-1705-4d05-aad9-e061034f6e0f · outbound
Acknowledging Focus Ambiguity in Visual Questions AmbigQA: Answering Ambiguous Open-domain Questions
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68e0eab0-52e4-4a67-ac79-734c16adcc21 · outbound
Acknowledging Focus Ambiguity in Visual Questions Gpt-4o system card, 2024
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 363dfbcc-f442-458a-8e2a-924e39ddd723 · outbound
Acknowledging Focus Ambiguity in Visual Questions Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8ac7ffbb-b468-47b6-a20f-8cf67585aa3d · outbound
Acknowledging Focus Ambiguity in Visual Questions Referring ex- pression comprehension: A survey of methods and datasets
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17a8c0ae-0a7f-49ab-9e1b-da757aca4039 · outbound
Acknowledging Focus Ambiguity in Visual Questions PACO: Parts and Attributes of Common Objects
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f89d040e-cef8-4fca-91b0-56061121aa2d · outbound
Acknowledging Focus Ambiguity in Visual Questions Glamm: Pixel grounding large multimodal model
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8703932d-f850-43e4-bead-c0062f26aa38 · outbound
Acknowledging Focus Ambiguity in Visual Questions Omnilabel: A challenging benchmark for language-based object detec- tion
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 89234697-5ee6-4e6b-a4a6-91a9535ebb22 · outbound
Acknowledging Focus Ambiguity in Visual Questions Towards vqa models that can read
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26e534b9-9523-48a3-87ec-8384c4fb2e41 · outbound
Acknowledging Focus Ambiguity in Visual Questions Why Did the Chicken Cross the Road? Rephrasing and Analyzing Ambiguous Questions in VQA
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16b95603-233b-4b4e-b1fd-20410e912d07 · outbound
Acknowledging Focus Ambiguity in Visual Questions Vizwiz- fewshot: Locating objects in images taken by people with visual impairments
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 81489aa4-755b-4d24-a8a7-cfc60c850f1f · outbound
Acknowledging Focus Ambiguity in Visual Questions Multimodal few-shot learning with frozen language models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f352b30-78d1-430d-b580-002d3179c707 · outbound
Acknowledging Focus Ambiguity in Visual Questions Modeling ambiguity, subjectivity, and diverging viewpoints in opinion question answering systems
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 56d6c376-c0e8-454a-ac1a-41bbf5954e86 · outbound
Acknowledging Focus Ambiguity in Visual Questions Twice opportunity knocks syn- tactic ambiguity: A visual question answering model with yes/no feedback
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 852d83ff-f24c-45f9-9e45-2afa099396b3 · outbound
Acknowledging Focus Ambiguity in Visual Questions Phrasecut: Language-based image segmen- tation in the wild
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f4dde751-9126-4443-b133-09ec9b1cb4a7 · outbound
Acknowledging Focus Ambiguity in Visual Questions Described object detection: Liberating ob- ject detection with flexible expressions
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5902d523-5713-4fa4-9500-9c8fb9168a77 · outbound
Acknowledging Focus Ambiguity in Visual Questions Visual Question Answer Diversity
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ee429aa1-a11f-433e-80bb-e867be02b912 · outbound
Acknowledging Focus Ambiguity in Visual Questions Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a186a13f-58b5-47e7-9f24-1900437161c8 · outbound
Acknowledging Focus Ambiguity in Visual Questions Visual7w: Grounded question answering in images
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e400f1b2-23e6-428f-ab94-f5cd58c9e832 · outbound
Acknowledging Focus Ambiguity in Visual Questions Object detection in 20 years: A survey.Proceed- ings of the IEEE, 111(3):257–276, 2023
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9d2d9f48-67ad-40c4-af68-19c80ac32b63 · inbound
GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning Acknowledging Focus Ambiguity in Visual Questions
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.