Pith. sign in

Paper Citation Record · LEDGER

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users

As of 17 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2509.06010.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06010 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:43:20.831117Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:50:19.280108Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:50:19.564824Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a71d3b19-4d05-44c9-bd8a-678cf942a4cf · outbound

This paper cites Vision-language model-based polyformer for recognizing visual ques- tions with multiple answer groundings.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Vision-language model-based polyformer for recognizing visual ques- tions with multiple answer groundings

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:23.704356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:18.629579Z digest=sha256:07addf4a8f99c55a3ad9a92aa07f6916e5ed3991471439e9ee0908fce462c2c1

Observation 7969617d-6b83-4407-b0c4-90192c9a1b68 · outbound

This paper cites Remote assistance for blind users in daily life: A survey about be my eyes.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Remote assistance for blind users in daily life: A survey about be my eyes

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:23.597271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:18.722461Z digest=sha256:52cc55825bc5b41cfcbabf5bd8c62c0b07de1e9f4b4a8e69b34e5b0eaf3c2943

Observation 038400d4-9bfc-4000-91bf-9a075d73091f · outbound

This paper cites Vqa therapy: Exploring answer differences by visually grounding answers.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Vqa therapy: Exploring answer differences by visually grounding answers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:23.481862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:18.812578Z digest=sha256:d88a87b7f68fdcd3381e55850baf7e6ca570c2e88a83e2266d9e410e7a15d886

Observation 88fb4172-b046-4b0a-97de-69bb7b728e07 · outbound

This paper cites Refining pseudo labeling via multi- granularity confidence alignment for unsupervised cross domain object detection.IEEE Transactions on Image Processing, 2025.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Refining pseudo labeling via multi- granularity confidence alignment for unsupervised cross domain object detection.IEEE Transactions on Image Processing, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:23.318942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:18.934653Z digest=sha256:4442840c3ab92f4ab7a958ab8aaf8d840ffeb94a6c8a3fba1e835410122d0d6f

Observation 17b92c3e-06ea-40ea-81ea-36dc0b1b2806 · outbound

This paper cites Vqask: a multimodal android gpt- based application to help blind users visualize pictures.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Vqask: a multimodal android gpt- based application to help blind users visualize pictures

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:23.155782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:19.079205Z digest=sha256:5551b23b5a70cf19bbc6fe8611c09975dd8a65891c364546fca32fdb5abf6d7d

Observation c35eb609-729a-4598-892a-f983135fa42f · outbound

This paper cites Towards understanding the use of mllm-enabled applications for visual interpretation by blind and low vision people.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Towards understanding the use of mllm-enabled applications for visual interpretation by blind and low vision people

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:23.013945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:19.198124Z digest=sha256:08c9d1f9ca3af67d03afc5d44897748bf12dc30adad38993019bb5003e408023

Observation 61ae3016-6ceb-4779-84ac-d3683cd7abef · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Vizwiz grand challenge: Answering visual questions from blind people

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:22.856931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:19.299838Z digest=sha256:6219d1a5e1b375c05d15981ae019442c2323a759ac7a4e2a9dacc5d94b620040

Observation 74eaca94-984d-4733-a645-be52a4ed3189 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Lora: Low-rank adaptation of large language models.ICLR, 1(2):3, 2022

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T04:43:19.427202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:43:19.427202Z digest=sha256:f5d42b55104853381ddd7ac4962df0f771beab1943f351b8ff959b5a0f593965

Observation 22e64f2a-a691-47bb-921b-5355012f4667 · outbound

This paper cites Consistency and uncertainty: Identifying unre- liable responses from black-box vision-language models for selective visual question answering.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Consistency and uncertainty: Identifying unre- liable responses from black-box vision-language models for selective visual question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:22.658833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:19.511234Z digest=sha256:3e81fd0543413e105fcdd48d3e57693606fae5a671ac9fa7ea68fc3149c89d46

Observation 3e5d52a3-e44c-4d03-858f-be4c7f564a44 · outbound

This paper cites Dual-branch fusion with style modulation for cross-domain few-shot semantic segmentation.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Dual-branch fusion with style modulation for cross-domain few-shot semantic segmentation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:22.540118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:19.621669Z digest=sha256:1ec1e8e267faa7c141651ada8e0887db6f1e9ec84f2da4dc20922a7132e47620

Observation 13112c59-87cd-4b19-bb27-628a24cc6a46 · outbound

This paper cites Natural language understanding and inference with mllm in visual question answering: A survey.ACM Computing Surveys, 57(8):1–36, 2025.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Natural language understanding and inference with mllm in visual question answering: A survey.ACM Computing Surveys, 57(8):1–36, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:22.407626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:19.745316Z digest=sha256:92a9b696a027071d6a22a8dedac316a0971a34552f50dace2808c6baba1c643d

Observation 8a0b82b8-71f9-4ab7-9d16-3712a3cb19fa · outbound

This paper cites Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:22.242938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:19.846109Z digest=sha256:44edc27e5f982df8cdbf25c88a71f96e870e7071436dcf0bf2d10818e9804381

Observation 59ab5756-76b5-4888-b221-fabfbf80582e · outbound

This paper cites Polyformer: Referring image segmentation as sequential polygon generation.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Polyformer: Referring image segmentation as sequential polygon generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:22.135019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:20.005306Z digest=sha256:8f24d30a0c45d55d156eaec016da4e20420d734ded712e502eefaf576d8f49d8

Observation 4cc3c410-d558-428a-a2d3-d7f83cd29085 · outbound

This paper cites An astute assistive device for mobility and object recognition for visually impaired people.IEEE Transactions on Human-Machine Systems, 49(5):449–460, 2019.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users An astute assistive device for mobility and object recognition for visually impaired people.IEEE Transactions on Human-Machine Systems, 49(5):449–460, 2019

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:22.046250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:20.210949Z digest=sha256:4552ccf5ad809acad31d8872aefd44bb335f8a54d69bacba1728382bbdade0c0

Observation 51efa02b-fb72-4e61-b604-ec675b7df3fa · outbound

This paper cites Dynamic conceptional con- trastive learning for generalized category discovery.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Dynamic conceptional con- trastive learning for generalized category discovery

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:21.954258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:20.357680Z digest=sha256:a9f72f50ccd5d04d568fdefeb6d9777d18440af66c0513de098266a41d2309b8

Observation 73ebf983-0b0a-4f26-9844-b897620e2f23 · outbound

This paper cites Advances in few-shot action recognition: A comprehensive review.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Advances in few-shot action recognition: A comprehensive review

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:21.873693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:20.466310Z digest=sha256:cd72e490f079c47e8225470664f705b651b1b15866cf3f51d0430185409e8cdb

Observation 2743f3ba-f3bf-4408-b218-19313375d57a · outbound

This paper cites DARE: Diverse Visual Question Answering with Robustness Evaluation.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users DARE: Diverse Visual Question Answering with Robustness Evaluation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-05T04:43:21.034136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:20.547086Z digest=sha256:476c7a6097cb7afb4a8ffab9a4da2723d8d02ac06ea1ff3462dcf39ffde78d38

Observation 7cfad7f6-621b-40f6-80f5-cb3990f11da5 · outbound

This paper cites the smart vision glasses.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users the smart vision glasses

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:21.753393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:20.632394Z digest=sha256:25e3ece5b94901f2d99dd173fdc71ef343128a456371ea02a891bfcd8242515e

Observation 55a58566-607d-42f3-8275-e293688f50dc · outbound

This paper cites A survey of 17 indoor travel assistance systems for blind and visually impaired people.IEEE Transactions on Human-Machine Systems, 52(1):134–148, 2021.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users A survey of 17 indoor travel assistance systems for blind and visually impaired people.IEEE Transactions on Human-Machine Systems, 52(1):134–148, 2021

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:21.627116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:20.706463Z digest=sha256:b9e8ce02311afe362214dc613974c899948ba1dfcf39979fc5e4e7f6a8af9e92

Observation 3793b5f3-9b50-4591-b94c-5f3bf8166f85 · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:21.422605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:20.778291Z digest=sha256:048495c4a65c9c0606b09911280967459ca036536c7c066551ce35ead4989467

Observation c3c93b60-2266-480f-a13a-eabe623cbcae · outbound

This paper cites A survey on vqa: Datasets and approaches.

BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users A survey on vqa: Datasets and approaches

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T04:43:21.238294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T04:43:20.831117Z digest=sha256:735a99a2768063985fd9c695674ce87d151ecf84c592c06ce8a17dd2e9c1fefd

Pith citing papers

Observation ed1391d9-df16-421e-b77a-79a1b2965b96 · inbound

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems cites this paper.

How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems BLaVe-CoT: Consistency-Aware Visual Question Answering for Blind and Low Vision Users

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-12T00:50:19.569534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T00:50:19.280108Z digest=sha256:e8e377cc97f7811de8c47b26b859a8bd6806ae78b4c2b3a9a424c4f4054080c7