Pith. sign in

Paper Citation Record · LEDGER

Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:1612.00837.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1612.00837 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:10:11.229978Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c44e8270-1ad1-49bd-a8b2-0d0bb52c55eb · inbound

Language Tasks and Language Games: On Methodology in Current Natural Language Processing Research cites this paper.

Language Tasks and Language Games: On Methodology in Current Natural Language Processing Research Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T10:38:39.034222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T10:38:39.034222Z digest=sha256:cbce739958276afebed0fffedebcb7fa45e67e318f6e60af82addf13f6c3a6e1

Observation 8f07b351-18dd-4c1d-ab18-1cc9f23a84a7 · inbound

Preserving Knowledge in Large Language Model with Model-Agnostic Self-Decompression cites this paper.

Preserving Knowledge in Large Language Model with Model-Agnostic Self-Decompression Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:13:39.503725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-24T00:09:52.093810Z digest=sha256:d53ffe7f2f14c43b68ef5e8a716ed30dd0331341b049c99d84967a7f07b8c838

Observation c0666ad8-9fba-4350-963b-0d05427ba75c · inbound

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability cites this paper.

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:04:52.066207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:04:52.066207Z digest=sha256:05b59aec35c88e5f57fe9db27a777c0413e82a985fc3f24dd48b199ba953cd12

Observation f59dcad5-8993-44fe-b3f2-142718843a77 · inbound

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering cites this paper.

FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:21.172080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:21.172080Z digest=sha256:1817e0979cb0dc6762cb3d2e712cdfa18c6bb3bda82679c4c5e66d58cbe11d9c

Observation 58a9dfed-de02-42d1-b30e-b7a62d530626 · inbound

Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities cites this paper.

Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:36:01.277024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:36:01.277024Z digest=sha256:1d4781dd403c36e2b373e1c2aa9c31c30793be7ca7b3e2bef429de16225f354a

Observation fa3ecfab-4878-4640-986d-eb42adafab7c · inbound

Differential Multimodal Transformers cites this paper.

Differential Multimodal Transformers Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:20.062426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:20.062426Z digest=sha256:a39bad0ae63415be02235cc772ed60d3015756cabf508a26aae7ce42626d092c

Observation a9bdf05d-29e1-4dd4-8334-a4e19eb2017e · inbound

PlantExpertVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science cites this paper.

PlantExpertVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:10:11.229978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:10:11.229978Z digest=sha256:3b3a6026b83d96b1587eefae8001e56b9600d789436d887a9dc29188d057bcf7

Observation 59b73ab2-1193-4a1f-ab8a-02923d8feecb · inbound

CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering cites this paper.

CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.200709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T07:09:48.239662Z digest=sha256:f6777ecf5d2693238a9625e0120fa470a6cfc2b650b7edd054f02e9c63e57fdb

Observation 25482ded-30fc-4dd8-9c94-f3f2d3eac16a · inbound

LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models cites this paper.

LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:15:50.368821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T05:11:01.039309Z digest=sha256:632c05d5a20b9e4d6e57c9aef3fb14127487b1b99f2cfabb52b88ec038655f77

Observation 25d1d983-259b-459b-bc68-c575850fe223 · inbound

EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving cites this paper.

EgoDyn-Bench: Evaluating Ego-Motion Understanding in Vision-Centric Foundation Models for Autonomous Driving Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:19:47.174538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T00:09:18.068337Z digest=sha256:d065a7323b6eff0822e1a3dc7c5dadb6f5f462a20a2e000307afb96db5a750fc

Observation a832c5d0-bdd4-47c6-94f1-2f74dcdc0bfa · inbound

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics cites this paper.

When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:36:26.509150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T10:51:38.455604Z digest=sha256:e2ec2f6b9d653177469fe74e73368ae210b7a925c2c03f475a708859f65e647f

Observation 6cc0774c-d953-4fd2-a1f2-000112688b05 · inbound

SPARC: A Multi-Agent System for Electrical Circuit Question Answering cites this paper.

SPARC: A Multi-Agent System for Electrical Circuit Question Answering Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T18:57:17.102609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:44:42.020128Z digest=sha256:41a00a3472729d109bf006c7aed70a0c8fa192215eb31e0bb83a363d187b107a

Observation fb7a23d7-1e3a-44c9-af28-11d6a880f651 · inbound

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference cites this paper.

SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-04T12:59:52.416002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T05:41:39.052865Z digest=sha256:b49d6e025b48546b70648a3ccf3455fefe3bd25be289a82b9e9c4ec00c0793f4

Observation a32ba699-9e1a-4bd8-9cbc-a51f9eeb0929 · inbound

ChatImage: Navigating Long-Form LLM Answers through Interactive Images cites this paper.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-07T19:34:06.385691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:3401ed2fd96894b319d91dc9ac02eafce9bc08c3d2a467c917bdb30a569d43f6

Observation 9a380d2e-0e92-4e99-ab94-937e0b96ea69 · inbound

Searching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language Models cites this paper.

Searching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language Models Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T19:13:26.665422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T19:13:26.665422Z digest=sha256:c40a8f3b4247532a225faf9bcc9788eea984e145d1c3e6173833c94a4dc0286c

Observation bc3fd0ed-4c5a-4351-870e-0fd3dbf05f7f · inbound

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models cites this paper.

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T16:41:28.885987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:41:28.885987Z digest=sha256:39743022076e5be5a045893fdf26f6c5b4b5f1d2d16e481e77f6c9acbe92d182

Observation 12a72b7e-cfe8-4a6d-a7f5-4f4b45bf6057 · inbound

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware cites this paper.

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:02.790931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:02.790931Z digest=sha256:9c23f4c060c833737001b7d186920128cc4669a95d8f9ba61a06337a86c55086

Observation 65c7b628-816a-453e-987e-e9bcfe67fc42 · inbound

TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs cites this paper.

TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T05:29:59.232774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:29:59.232774Z digest=sha256:1a210e8371741e9ee3ee4d5978879f31615ba671b454bd5c1fa7b441239aa441

Observation 794518a2-b532-4a00-8c5c-7777234decef · inbound

TomaMMU: A Comprehensive Multimodal Understanding Benchmark for Tomato Leaf Diseases cites this paper.

TomaMMU: A Comprehensive Multimodal Understanding Benchmark for Tomato Leaf Diseases Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:31:22.208670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:31:22.208670Z digest=sha256:e0b228ec43ec3cbb6995c1f3efc819cd55d0829d2173d096d150802d9b613a39

Observation bbaafca0-6a2d-4307-b0b9-f085588084db · inbound

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL cites this paper.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.428086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.428086Z digest=sha256:e1579f994ee9034ec2f25f3736763d67751badd6865e8fc778e1073b30d976a9