Pith. sign in

Paper Citation Record · LEDGER

Acknowledging Focus Ambiguity in Visual Questions

As of 21 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2501.02201.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02201 v2

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:19:13.746242Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T02:03:52.566413Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T04:00:54.597918Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1db03ffc-a8cd-4cb3-95ea-be41ea326195 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Acknowledging Focus Ambiguity in Visual Questions Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.539807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.539807Z digest=sha256:9e68a0d309549cd1e20b0fd1e357c6f0ff0e0444f514cd0c37665de1dbe5293d

Observation abb0115e-9ccc-42c3-a432-c218d43d2764 · outbound

This paper cites Vqa: Visual question answering.

Acknowledging Focus Ambiguity in Visual Questions Vqa: Visual question answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.544017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.544017Z digest=sha256:d214a8e6a6eddec0179943be24b3720e5577050cfdc7484bea61a4495f73b412

Observation 046961b5-25f2-4b06-adae-f8c25fdb0829 · outbound

This paper cites Qwen2.5-VL Technical Report.

Acknowledging Focus Ambiguity in Visual Questions Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.547498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.547498Z digest=sha256:3165cd0c100d644254c2a1408822a0372b5ea472de253eaf63a1f881385bb4a4

Observation 2caccec5-8c0d-4c12-985c-10cf576219e7 · outbound

This paper cites Why does a visual question have different answers? In Proceedings of the IEEE International Conference on Computer Vision , pages 4271–4280, 2019.

Acknowledging Focus Ambiguity in Visual Questions Why does a visual question have different answers? In Proceedings of the IEEE International Conference on Computer Vision , pages 4271–4280, 2019

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.382252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.551497Z digest=sha256:0ea1908569169d3ef38650258232e5775984de599a9aa297fc23da6185f5123b

Observation 35cf7e17-5a7b-4327-bf6b-7dbb48f71430 · outbound

This paper cites Grounding answers for visual questions asked by visually impaired people.

Acknowledging Focus Ambiguity in Visual Questions Grounding answers for visual questions asked by visually impaired people

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.369517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.555292Z digest=sha256:7784356e3f87d808703f1b892d0d8c4f57d456be277badc3a2c9e4b0eb6a9e51

Observation 5e6765d8-826e-47b3-9dcc-b2f0c189e2f7 · outbound

This paper cites Vqa therapy: Exploring answer differences by visually ground- ing answers.

Acknowledging Focus Ambiguity in Visual Questions Vqa therapy: Exploring answer differences by visually ground- ing answers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.357473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.558672Z digest=sha256:e113305912aa4c6ac29de41d17c79dcef7b640db4191a022a4686dd6c93f178a

Observation 635550d4-1e76-413f-a77b-8615a1694fc8 · outbound

This paper cites Fully authentic visual question answering dataset from online communities.

Acknowledging Focus Ambiguity in Visual Questions Fully authentic visual question answering dataset from online communities

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.344905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.562069Z digest=sha256:bf6dae4740d2b58e2e1b84b6574470e05eb82c7b06d0b502bea1c2a24c0c0966

Observation 011efa55-1d4e-4a66-ad1c-9fc2c1d88cf9 · outbound

This paper cites Cops-ref: A new dataset and task on composi- tional referring expression comprehension.

Acknowledging Focus Ambiguity in Visual Questions Cops-ref: A new dataset and task on composi- tional referring expression comprehension

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.333153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.565125Z digest=sha256:ff3b0d82f4e887777d32e2bd039967cf902daf62a3e49afbe1778c8b995abbd2

Observation a0e3da99-d8f5-4681-a206-039167b74e9e · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Acknowledging Focus Ambiguity in Visual Questions How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.568825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.568825Z digest=sha256:57c7d74c0f91270a93efca0466df4732da3407091809e80d42c023d36aac3980

Observation 9db10d44-fa89-48d2-a68f-fa6e216b624f · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Acknowledging Focus Ambiguity in Visual Questions Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.573077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.573077Z digest=sha256:ec029af87d710883a388e8ac6ab7609a8f420ed564d456aa424ef902e49a7327

Observation 3a3bbdbb-24c3-47b5-93be-e98729387ccb · outbound

This paper cites Resolving Language and Vision Ambiguities Together: Joint Segmentation & Prepositional Attachment Resolution in Captioned Scenes.

Acknowledging Focus Ambiguity in Visual Questions Resolving Language and Vision Ambiguities Together: Joint Segmentation & Prepositional Attachment Resolution in Captioned Scenes

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:19:13.895356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.577309Z digest=sha256:058e1c33ac9f47028dcf320860863b6815cf57fdbb5466f01ea37def0822ade3

Observation b3bcbc7a-8ea3-4ab5-a549-b2e4cfcff3b4 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Acknowledging Focus Ambiguity in Visual Questions Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.581827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.581827Z digest=sha256:fd47c97b545c1a3c61db41ea2b210321a6b92f93f4d2c9c74eabcaa36ed99b4f

Observation b1e47f0a-baf4-4de6-b1bf-343d711dfa2c · outbound

This paper cites Zero-shot and few-shot video question answering with multi-modal prompts.

Acknowledging Focus Ambiguity in Visual Questions Zero-shot and few-shot video question answering with multi-modal prompts

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.321971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.586224Z digest=sha256:3a5c68ce673dba1f2ce998bcc2829a7f44363dd6a5f73a22971fd03f91ee51bd

Observation c48e1e52-e66b-416e-9b4d-6f96680ec71c · outbound

This paper cites Vqs: Linking segmentations to questions and answers for supervised attention in vqa and question-focused semantic segmentation.

Acknowledging Focus Ambiguity in Visual Questions Vqs: Linking segmentations to questions and answers for supervised attention in vqa and question-focused semantic segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.311226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.590122Z digest=sha256:27a5a468a77d5cf2eb914df4e64bf4e3e1f6c3fc65ca60b6211b3a90c95fa9e1

Observation 853b4b89-a37d-4f41-8f87-377224909b5a · outbound

This paper cites Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering.

Acknowledging Focus Ambiguity in Visual Questions Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.300267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.593974Z digest=sha256:7b3ec0e5a42767518029592dcfefc73a7198ee34f29fd8276192035fc39939b0

Observation 34991dcb-c056-4db1-b3ab-d884437dda1a · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Acknowledging Focus Ambiguity in Visual Questions Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.597704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.597704Z digest=sha256:ba1508f2bb848b669ad5377519cb53531e12d73553aff2d00bb845a6f672de98

Observation 57477597-2063-4653-83f7-ad35ea92d844 · outbound

This paper cites Abg-coqa: Clarifying ambiguity in conversa- tional question answering.

Acknowledging Focus Ambiguity in Visual Questions Abg-coqa: Clarifying ambiguity in conversa- tional question answering

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.279047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.601696Z digest=sha256:c9399f8f79ed5fb8fc2fe8e64ad74b4aa2cc1bd9c79b190612d54af8d0e1b513

Observation 57820aaf-2d32-4f42-b173-60d5980fc62f · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Acknowledging Focus Ambiguity in Visual Questions Lvis: A dataset for large vocabulary instance segmentation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.266509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.605590Z digest=sha256:04094e3f203efb6da75d4a77d81af698c96ab334310b9658641115f0fc730f95

Observation 4201e5f5-1fef-4fd7-9c2b-d9caf787616d · outbound

This paper cites Crowdverge: Predicting if people will agree on the answer to a visual question.

Acknowledging Focus Ambiguity in Visual Questions Crowdverge: Predicting if people will agree on the answer to a visual question

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.253575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.610550Z digest=sha256:66fed46c5dc176d3ac4ab97bf9bb51d05894f40ce7631ecdf53c099abe2eb924

Observation 2e105630-43e1-464d-b41e-f46e7fb7c367 · outbound

This paper cites Gurari, K.

Acknowledging Focus Ambiguity in Visual Questions Gurari, K

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.241375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.614613Z digest=sha256:9acd4516e63e4297b5fa94769bb7fa3b2c53d8c14102fc9178626e278366d7fe

Observation ea230c51-3467-41b8-9204-36ecdbb20ccc · outbound

This paper cites Predicting foreground object ambiguity and efficiently crowdsourcing the segmentation (s).

Acknowledging Focus Ambiguity in Visual Questions Predicting foreground object ambiguity and efficiently crowdsourcing the segmentation (s)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.230175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.619382Z digest=sha256:6aac945a93aea8036fa3bfa5aa265795696a0730675739f7edd71a037bb34951

Observation c4a3b52b-ac85-46c4-8ffb-64246d111d63 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Acknowledging Focus Ambiguity in Visual Questions Vizwiz grand challenge: Answering visual questions from blind people

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.623315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.623315Z digest=sha256:d6e8235e52269ff558cbd1e5a43c5e61de69337f404528de9129302a2be40e87

Observation 4804eaf0-a390-47bf-a4f2-ea42c7e351c2 · outbound

This paper cites A survey on instance segmentation: state of the art.International jour- nal of multimedia information retrieval, 9(3):171–189, 2020.

Acknowledging Focus Ambiguity in Visual Questions A survey on instance segmentation: state of the art.International jour- nal of multimedia information retrieval, 9(3):171–189, 2020

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.210965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.627755Z digest=sha256:feed4d27fea5c6f625d8b8ac813296709f160b1037e3fab3f91b05921a7b9c30

Observation 8a8c6159-872b-4490-b94c-f1783ee6edfa · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Acknowledging Focus Ambiguity in Visual Questions Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.197739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.631499Z digest=sha256:2b22d6847fa2ee6ca93aec051088cc6a27af13e130d1249b75c1427cfd65d7d6

Observation a8dfdd65-a7c9-4e4a-b377-45a87af2f0c7 · outbound

This paper cites Long-Form Answers to Visual Questions from Blind and Low Vision People.

Acknowledging Focus Ambiguity in Visual Questions Long-Form Answers to Visual Questions from Blind and Low Vision People

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.635184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.635184Z digest=sha256:5272d27f9ab0cd97e141f70a939e572f6d2f6d379312cdc01492a254372ffd05

Observation 850fb64c-d769-4a76-ad0b-9fea91b5cc17 · outbound

This paper cites Salient object detection: A discriminative regional feature integration approach.

Acknowledging Focus Ambiguity in Visual Questions Salient object detection: A discriminative regional feature integration approach

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.185954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.639245Z digest=sha256:412b096cca6afa456d34b70317b29fc707d9cb83df4b69517a6b53182c5139e0

Observation 46eb0406-f91b-444f-8b4c-68b9d3653621 · outbound

This paper cites Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models.

Acknowledging Focus Ambiguity in Visual Questions Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.642999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.642999Z digest=sha256:d65acbd8eb0e0a9af12b7468825ee34fed48ff9e335f1dc9861d5554bdda056f

Observation 3c97303a-e98f-4b4f-9231-252b696bc887 · outbound

This paper cites Segment any- thing.

Acknowledging Focus Ambiguity in Visual Questions Segment any- thing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.173503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.646478Z digest=sha256:daef2001f7299aaf0dcae1f57a5164aaebd7ba2b5fa5ae0ccabdb56b92d12610

Observation 4f380e9f-5f23-40d4-8959-86a1e95516b4 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Acknowledging Focus Ambiguity in Visual Questions Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.161421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.649803Z digest=sha256:2ab6315b657d886ef8f00b29eb069d2f89fbd34f1a86abab4a2d78d742f91283

Observation 496ad413-1645-422f-8535-747f1edabc8f · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Acknowledging Focus Ambiguity in Visual Questions SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.652557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.652557Z digest=sha256:5818e842eef75c990034c495d2b63dc9767635317cef51042f1313881fbc1795

Observation 6195d37c-5db3-4e36-a8d1-09fe210799ee · outbound

This paper cites Microsoft coco: Common objects in context.

Acknowledging Focus Ambiguity in Visual Questions Microsoft coco: Common objects in context

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.655515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.655515Z digest=sha256:a793ca8fddcc13bd8f7d84330f8f45a8b1a2a9f67e32457e2d3bb0c6f8d85bf0

Observation 214bc35a-d889-44c9-9e38-7e83f092d068 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, January 2024.

Acknowledging Focus Ambiguity in Visual Questions Llava-next: Im- proved reasoning, ocr, and world knowledge, January 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.142335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.658380Z digest=sha256:a83463b38e1ac107ecddc8c684d048f5d3cf50f6e286279e8dd1dbf1e444df88

Observation 0905dfff-9a34-430b-85ea-f7fd155a9d64 · outbound

This paper cites Learning to detect a salient object.

Acknowledging Focus Ambiguity in Visual Questions Learning to detect a salient object

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.129885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.661432Z digest=sha256:339fe5360b9bde4ae8eefe56d2507854876b8c0f2ef76363dfa835698bdebaff

Observation 34b8eedb-3687-4a2e-be56-31c07a7ade74 · outbound

This paper cites Storytelling with image data: a systematic review and comparative anal- ysis of methods and tools.

Acknowledging Focus Ambiguity in Visual Questions Storytelling with image data: a systematic review and comparative anal- ysis of methods and tools

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.119089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.664520Z digest=sha256:33be2031dbb7a818ec9c041076871817a51f79041b8320062ba06d68cea88c42

Observation efe7ec67-707a-43b3-a622-7daad9b8fd6c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Acknowledging Focus Ambiguity in Visual Questions MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.667579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.667579Z digest=sha256:c98534d81845af8dbbf57c48193e9f56ace511f6372c4032acbdee43a932a4ad

Observation 62494ba0-6807-4115-b024-579d2c88284d · outbound

This paper cites an unresolved cited work.

Acknowledging Focus Ambiguity in Visual Questions Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:14.107800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.670906Z digest=sha256:2f0dad7b6facb426aa0b14e71a696a113da338bafc3d0892ee62d6188f78bd20

Observation 90ee9174-dae7-4426-a98c-5ab9c659b7c8 · outbound

This paper cites Resolving ambi- guities in text-to-image generative models.

Acknowledging Focus Ambiguity in Visual Questions Resolving ambi- guities in text-to-image generative models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.096850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.674226Z digest=sha256:541b3b23d61ecbfbba4ae8f37990a78f37414a38a50cae663fb3d9619f5aac4d

Observation 08c67028-1705-4d05-aad9-e061034f6e0f · outbound

This paper cites AmbigQA: Answering Ambiguous Open-domain Questions.

Acknowledging Focus Ambiguity in Visual Questions AmbigQA: Answering Ambiguous Open-domain Questions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.677458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.677458Z digest=sha256:ddf679db9390b6cc2ccc72d7c44d0b4ca6134bd654cbb8701ca1850b92d286e5

Observation 68e0eab0-52e4-4a67-ac79-734c16adcc21 · outbound

This paper cites Gpt-4o system card, 2024.

Acknowledging Focus Ambiguity in Visual Questions Gpt-4o system card, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.084045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.680844Z digest=sha256:741ba4178ad5a129c3d4e5e122b45ca7724b028be75af305f00acde31d933399

Observation 363dfbcc-f442-458a-8e2a-924e39ddd723 · outbound

This paper cites Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models.

Acknowledging Focus Ambiguity in Visual Questions Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:19:13.812809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.684378Z digest=sha256:213540b17b665ab2236288ba6904157155d30bf8ccbf0498dba5de7a6f217ec8

Observation 8ac7ffbb-b468-47b6-a20f-8cf67585aa3d · outbound

This paper cites Referring ex- pression comprehension: A survey of methods and datasets.

Acknowledging Focus Ambiguity in Visual Questions Referring ex- pression comprehension: A survey of methods and datasets

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.687703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.687703Z digest=sha256:337b1529ddd11fe88b05c5b1ec472f5e17f718515e9d4dc7b67ffd672564cc26

Observation 17a8c0ae-0a7f-49ab-9e1b-da757aca4039 · outbound

This paper cites PACO: Parts and Attributes of Common Objects.

Acknowledging Focus Ambiguity in Visual Questions PACO: Parts and Attributes of Common Objects

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.690739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.690739Z digest=sha256:f8687a01fb0b38ffe90719a313de75b79767f653c215bac6dce488682c703a42

Observation f89d040e-cef8-4fca-91b0-56061121aa2d · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Acknowledging Focus Ambiguity in Visual Questions Glamm: Pixel grounding large multimodal model

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.064636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.694329Z digest=sha256:91e9096411ea55ea26962a39755385af7f09688c18030b50aec09fd3e4ba340a

Observation 8703932d-f850-43e4-bead-c0062f26aa38 · outbound

This paper cites Omnilabel: A challenging benchmark for language-based object detec- tion.

Acknowledging Focus Ambiguity in Visual Questions Omnilabel: A challenging benchmark for language-based object detec- tion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.051910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.697717Z digest=sha256:6859f7cf2950f5e17158424d08dcec43ddc03da8a92f7671d570dd927c0ee022

Observation 89234697-5ee6-4e6b-a4a6-91a9535ebb22 · outbound

This paper cites Towards vqa models that can read.

Acknowledging Focus Ambiguity in Visual Questions Towards vqa models that can read

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.701194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.701194Z digest=sha256:3ec5c0d5512e32dce8cde77459572fbea022e51c3ffa75218b89773d5e2d61c5

Observation 26e534b9-9523-48a3-87ec-8384c4fb2e41 · outbound

This paper cites Why Did the Chicken Cross the Road? Rephrasing and Analyzing Ambiguous Questions in VQA.

Acknowledging Focus Ambiguity in Visual Questions Why Did the Chicken Cross the Road? Rephrasing and Analyzing Ambiguous Questions in VQA

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.705234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.705234Z digest=sha256:17498946a569f1250fe6c58de3f03e1eb81bec568b34945595e1358a1ed11740

Observation 16b95603-233b-4b4e-b1fd-20410e912d07 · outbound

This paper cites Vizwiz- fewshot: Locating objects in images taken by people with visual impairments.

Acknowledging Focus Ambiguity in Visual Questions Vizwiz- fewshot: Locating objects in images taken by people with visual impairments

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.033908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.709388Z digest=sha256:1a48c814c53ab2fca60efb2672db606e11e0a2fde0344a4ca84aca61c6e396ee

Observation 81489aa4-755b-4d24-a8a7-cfc60c850f1f · outbound

This paper cites Multimodal few-shot learning with frozen language models.

Acknowledging Focus Ambiguity in Visual Questions Multimodal few-shot learning with frozen language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:13.713405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:13.713405Z digest=sha256:47879a54513437b93a5415a76be658aa1f6c9f129c72df4c272b0394b2204534

Observation 4f352b30-78d1-430d-b580-002d3179c707 · outbound

This paper cites Modeling ambiguity, subjectivity, and diverging viewpoints in opinion question answering systems.

Acknowledging Focus Ambiguity in Visual Questions Modeling ambiguity, subjectivity, and diverging viewpoints in opinion question answering systems

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.014872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.717524Z digest=sha256:b4ee865b0d3a4c433fd5312142c7599edd05f73bdad0277718f2849bc383e3d9

Observation 56d6c376-c0e8-454a-ac1a-41bbf5954e86 · outbound

This paper cites Twice opportunity knocks syn- tactic ambiguity: A visual question answering model with yes/no feedback.

Acknowledging Focus Ambiguity in Visual Questions Twice opportunity knocks syn- tactic ambiguity: A visual question answering model with yes/no feedback

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:14.003034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.721701Z digest=sha256:ae9f8557116448b1a7b14e9927ffe9731de5159089d7864932b153e43ca5cf3a

Observation 852d83ff-f24c-45f9-9e45-2afa099396b3 · outbound

This paper cites Phrasecut: Language-based image segmen- tation in the wild.

Acknowledging Focus Ambiguity in Visual Questions Phrasecut: Language-based image segmen- tation in the wild

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:13.990856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.725793Z digest=sha256:e0fd43c606bb6e88cda54a243c3974a563a148851a3af724dc9065bcc1f29731

Observation f4dde751-9126-4443-b133-09ec9b1cb4a7 · outbound

This paper cites Described object detection: Liberating ob- ject detection with flexible expressions.

Acknowledging Focus Ambiguity in Visual Questions Described object detection: Liberating ob- ject detection with flexible expressions

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:13.979253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.729933Z digest=sha256:d30ce3c5001eb3a134fba604731faaab21eea738305914cf9ff2074eb25b9224

Observation 5902d523-5713-4fa4-9500-9c8fb9168a77 · outbound

This paper cites Visual Question Answer Diversity.

Acknowledging Focus Ambiguity in Visual Questions Visual Question Answer Diversity

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:13.968710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.734248Z digest=sha256:84bbbf4f10b7d00c1c9b644016028a69febbf76d053af2c8aa64eb6d2237daad

Observation ee429aa1-a11f-433e-80bb-e867be02b912 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

Acknowledging Focus Ambiguity in Visual Questions Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:13.958709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.737940Z digest=sha256:7851acac091604e4155744bc582c2f6355ef2208360a59daab1069668b81e6c9

Observation a186a13f-58b5-47e7-9f24-1900437161c8 · outbound

This paper cites Visual7w: Grounded question answering in images.

Acknowledging Focus Ambiguity in Visual Questions Visual7w: Grounded question answering in images

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:13.948491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.741831Z digest=sha256:faa238562d613dce4765b3237074a2ea51e80f98628ae5c893ffb188b1959939

Observation e400f1b2-23e6-428f-ab94-f5cd58c9e832 · outbound

This paper cites Object detection in 20 years: A survey.Proceed- ings of the IEEE, 111(3):257–276, 2023.

Acknowledging Focus Ambiguity in Visual Questions Object detection in 20 years: A survey.Proceed- ings of the IEEE, 111(3):257–276, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:13.938505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T22:19:13.746242Z digest=sha256:14718948574bbe613a2c550a8a10c607a5c40a6185eb7a90258ad86464b6ec85

Pith citing papers

Observation 9d2d9f48-67ad-40c4-af68-19c80ac32b63 · inbound

GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning cites this paper.

GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning Acknowledging Focus Ambiguity in Visual Questions

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:54.599999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T02:03:52.566413Z digest=sha256:421f24763036c56599b237c1556fabda997b052f37562821224071c75ce01b0f