Pith. sign in

Paper Citation Record · LEDGER

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework

As of 20 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2502.00711.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00711 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:03:53.348354Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:33:44.595993Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T14:33:46.092736Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact6
  • verified fuzzy30
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 649cd177-51ed-4f04-9e08-d8538d083313 · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.108911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.108911Z digest=sha256:5c686f4529bb6be2b1a7d0bb06c2c4d8eec385acbd366975ac954acacaa25020

Observation d1e22241-9f19-4b57-a981-2e8df316fd40 · outbound

This paper cites Visual instruction tuning,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Visual instruction tuning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.114192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.114192Z digest=sha256:10827091214536f2b83c372b524af83fbb82c7569b3ea5c424742db498f97bda

Observation ddb8f5b2-24eb-467b-9b82-dc9abbc70039 · outbound

This paper cites Mdetr-modulated detection for end-to-end multi-modal understanding,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Mdetr-modulated detection for end-to-end multi-modal understanding,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.273883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.118804Z digest=sha256:d758ce0fe3ba204893a340ea00077eeaa0e5e071b82aa371e7acaee73eae8cfa

Observation cad07e75-4fab-4b70-b696-a2df3173f6ad · outbound

This paper cites Omni-smola: Boosting generalist multimodal models with soft mixture of low-rank experts,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Omni-smola: Boosting generalist multimodal models with soft mixture of low-rank experts,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.259335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.123583Z digest=sha256:c5994246228764083c029ad58958b0e3d6ede85d39eee9af181980f7fbd97582

Observation 78d15b40-0858-4e91-b723-e66af2381bdf · outbound

This paper cites Cola: A benchmark for compositional text-to-image retrieval,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Cola: A benchmark for compositional text-to-image retrieval,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.244820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.128033Z digest=sha256:2ede1ebb9cc5cb0d80d0602aa6bf31c4197d2c066d593de7ed65faa194d8651a

Observation fd748ad8-7833-406c-b731-c2e9de31d20c · outbound

This paper cites Visual programming: Compositional visual reasoning without training,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Visual programming: Compositional visual reasoning without training,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.230148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.132829Z digest=sha256:a794fcc4cec370e470d9c86ddd758435c2d255d06af128051dd69b4340f71fcc

Observation 13ca5d95-21ba-4f0c-83a7-a6ef98b275c2 · outbound

This paper cites Interpretable visual reasoning: A survey,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Interpretable visual reasoning: A survey,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.216095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.138224Z digest=sha256:e05b07bbd1693eb6b6b2c2556328204d99b98b3e14a78ff25d689857f569d926

Observation 9d897cae-9bd3-4bce-925a-d4d70a53e570 · outbound

This paper cites Rapper: Reinforced rationale-prompted paradigm for natural language explanation in visual question answering,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Rapper: Reinforced rationale-prompted paradigm for natural language explanation in visual question answering,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.200839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.143782Z digest=sha256:89a429b1302c7fc12e1bd81597d0425afc8b5c309d30066f4cfdf9af3b8e6217

Observation 87d224cb-6bc4-47bf-90d1-cea4ba5904f5 · outbound

This paper cites Rephrase, augment, reason: Visual grounding of questions for vision-language models,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Rephrase, augment, reason: Visual grounding of questions for vision-language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.183457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.148002Z digest=sha256:9e007fceada108a131f60a4df08766dcc235410f74131131ffe5f8814623524d

Observation 9c97c9b4-04b5-43ad-876c-31afd58a9186 · outbound

This paper cites Toward multi-granularity decision- making: Explicit visual reasoning with hierarchical knowledge,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Toward multi-granularity decision- making: Explicit visual reasoning with hierarchical knowledge,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.152331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.157639Z digest=sha256:75d27a60d56637d248d2353df852cf96e8d394885969ce7103c8179eb22718eb

Observation e6cc5687-cb05-4ac4-89df-ad9c12f108cb · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.162330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.162330Z digest=sha256:dbd43621303ebba14bdef3802bbd685c7b0fbf52cdd66090fb22d0a5316a8cb7

Observation 4ecd5353-0806-429f-bd7d-adc11a0573bb · outbound

This paper cites Vision–language model for visual question answering in medical imagery,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Vision–language model for visual question answering in medical imagery,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.137221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.167035Z digest=sha256:fb0e4f2fb4086cd7f35d0c08c4d823f2fb960ae48a28a802cd18eb693a9e0cf3

Observation 09bf68a5-9800-46fb-a189-f9a5aa0935ad · outbound

This paper cites Lingoqa: Visual question answering for autonomous driving,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Lingoqa: Visual question answering for autonomous driving,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.120736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.172078Z digest=sha256:69683ee22d3d566fb67336e4fde65e13ad2f83e150b6c5d5354a5fa2530a66ba

Observation 493c5b92-5199-4a6a-9220-f8dc3ab48917 · outbound

This paper cites VQA and Visual Reasoning: An Overview of Recent Datasets, Methods and Challenges.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework VQA and Visual Reasoning: An Overview of Recent Datasets, Methods and Challenges

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.176903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.176903Z digest=sha256:961347afa6dda0c03cbbe93816682d5e46d1ae3029948192841410225cd995b9

Observation a8ac84f1-acd6-40fd-a085-f22403b49a64 · outbound

This paper cites Dealing with Semantic Underspecification in Multimodal NLP.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Dealing with Semantic Underspecification in Multimodal NLP

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.181943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.181943Z digest=sha256:ef85c4202f7f4707e1dfa77639ae75c4c18f2a5a86b5dd8ab8f9446e30b076ab

Observation 201d67f5-58d4-40b1-87a2-6cb46d878e3b · outbound

This paper cites Open visual knowledge extraction via relation-oriented multimodality model prompting,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Open visual knowledge extraction via relation-oriented multimodality model prompting,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.103734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.187050Z digest=sha256:817e2713b353575457458b1460fda15e8c7fe32189a9445461e67327d3549220

Observation 45b14339-5b4d-41df-aaf5-dcf88d0e2447 · outbound

This paper cites PV2TEA: Patching Visual Modality to Textual-Established Information Extraction.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework PV2TEA: Patching Visual Modality to Textual-Established Information Extraction

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.675681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.196382Z digest=sha256:1640c3fa043bfe0c30c7d67e045a619c5696d296849f6e8b645143398646a6cd

Observation c41710f0-af35-4fab-8584-779222491bc2 · outbound

This paper cites Recurrent fusion network for image captioning,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Recurrent fusion network for image captioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.073290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.201585Z digest=sha256:db27e6b62b3fffc94e82cf377e05f6dea5abe0c4370d1ea6ce2550712d566662

Observation 9cc4f3b7-a6f7-44f6-a865-61ff0b598e07 · outbound

This paper cites Boosting image captioning with attributes,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Boosting image captioning with attributes,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.057716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.206094Z digest=sha256:1c0ea09cb84011cb245aa3c58e4e9ecb7757a3e692357dece334249369f268d0

Observation ec9a6f63-0b01-4463-8b84-a56fea172bf6 · outbound

This paper cites Promptcap: Prompt-guided image captioning for vqa with gpt-3,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Promptcap: Prompt-guided image captioning for vqa with gpt-3,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.042647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.210754Z digest=sha256:a7413f1ddd81a7072d9a37f1383b076ad1f847d9e4ec37d1f1f072d17ea8a109

Observation 2b3de275-8ce0-4c20-9023-989fc64a8130 · outbound

This paper cites Injecting semantic concepts into end-to-end image captioning,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Injecting semantic concepts into end-to-end image captioning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.026025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.215491Z digest=sha256:d4d9900121615b490c0d4ebf8411ea5765dca01fe0da9074ae20db288850fc6a

Observation 44a583f8-05d8-490e-9b87-d344108e3e9b · outbound

This paper cites Visual Commonsense based Heterogeneous Graph Contrastive Learning.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Visual Commonsense based Heterogeneous Graph Contrastive Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.652545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.220310Z digest=sha256:bbf763b9ddb0d9c466bd66a614fe94eb7c66d8e7d6c2e0864a5ac198fd8fd06b

Observation 576ee7d7-bb2a-411f-bda0-bc7818ca20da · outbound

This paper cites Covlm: Composing visual entities and relationships in large language models via communicative decoding,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Covlm: Composing visual entities and relationships in large language models via communicative decoding,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.011138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.225287Z digest=sha256:877f9f5e799815432561cf4254a6abb0dac2a319cd82280c092809c400a5abaa

Observation 859becd4-8893-49f5-8051-0a6a55c14891 · outbound

This paper cites Bridging knowledge graphs to generate scene graphs,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Bridging knowledge graphs to generate scene graphs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.995215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.229765Z digest=sha256:8b1eaa970cf7c6395a105ab45228feed4b094798eaaa5cc13be147d356be50e4

Observation 17da6dd9-5f85-4145-b893-db9cff8abcfb · outbound

This paper cites Co-training improves prompt-based learning for large language models,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Co-training improves prompt-based learning for large language models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.979908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.234699Z digest=sha256:6ab8adf3181bd8b4899fb8d61a374b426ae112783bdddd7b48a770500ed41100

Observation 352cb6f8-0e86-4031-9d49-c0d03e1ee2c6 · outbound

This paper cites Large language model as attributed training data generator: A tale of diversity and bias,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Large language model as attributed training data generator: A tale of diversity and bias,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.964531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.239804Z digest=sha256:a7a474586519660cd2eb5c463f33682e56debf095be2b240ef8afd36a2424218

Observation f23dd74d-2d9e-4d04-965a-e696057e3844 · outbound

This paper cites Large language models are zero-shot reasoners,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Large language models are zero-shot reasoners,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.244544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.244544Z digest=sha256:126568f50c5e31d2d7bfd9acd4af658891e69dd30393e421d5ee63dd9b392187

Observation 72faaa99-7ef1-4082-8c89-e4dd2dface47 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.249363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.249363Z digest=sha256:acd16c22c01607154b47d257dcfa9a9e423dad5d3c4e51e1fc216ca94d23b706

Observation 3bd099ee-9f1a-4352-b2f0-45ae8faa3e49 · outbound

This paper cites e-vil: A dataset and benchmark for natural language explanations in vision-language tasks,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework e-vil: A dataset and benchmark for natural language explanations in vision-language tasks,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.937640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.254268Z digest=sha256:390ba5f9cf26dd66fa0e8b874dc31c1da00a913ada434f08a5c4c4009ea80343

Observation 37485df7-ccff-4907-a53b-d104118fabc5 · outbound

This paper cites NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.411763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.258591Z digest=sha256:c56db6ed86d97eb77a0ab3d4225e58608d12fa4e33c8632ed4f29d3501aa3267

Observation 7e32c95e-1594-4df8-a12c-5606734b071c · outbound

This paper cites Improving vision-and-language reasoning via spatial relations modeling,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Improving vision-and-language reasoning via spatial relations modeling,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.088306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.262983Z digest=sha256:a1889c5b0865976586631cb0d99eb9a875d5a29df378958e41a408c9724a29ee

Observation 581bbdad-363a-4dde-a6dd-c3153a860408 · outbound

This paper cites GPT-4 Technical Report.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework GPT-4 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.267186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.267186Z digest=sha256:ed7eac85c25f06206c1be2bcf40b904cd56088d4a8c0045961f502a1252c2678

Observation 49a49c73-976b-48ed-b03b-edf5ca7db438 · outbound

This paper cites Chain of thought prompting elicits reasoning in large language models,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Chain of thought prompting elicits reasoning in large language models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.920798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.271528Z digest=sha256:0b1d8364c741c50413e16e466f1d6bbbb9a42f89748486fd5f97d627f33af650

Observation 6301d550-8086-41db-9af7-2e5291357021 · outbound

This paper cites Re- flexion: Language agents with verbal reinforcement learning,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Re- flexion: Language agents with verbal reinforcement learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.276014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.276014Z digest=sha256:a3c3b35f36439732978eec429403d2baa660af6e9c0528d3f4744e2dd453e4e8

Observation 48cb50bd-3dce-4d0a-b28a-5375e5202022 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.280509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.280509Z digest=sha256:893d1ef29802aa0a1e3f8ec722f31ac6584ebf616913c251d610f698f94e19f7

Observation ffc1b197-7c6f-4a7c-ab72-dbc64b4aa10b · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework A-okvqa: A benchmark for visual question answering using world knowledge,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.883686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.284831Z digest=sha256:8c994f2caf0d97aec2c3d878d79acd4d4687cd713f256a527bc93441e28db10c

Observation 2c7d8b91-5454-4649-8acd-ffe132606bac · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Vizwiz grand challenge: Answering visual questions from blind people,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.866801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.288845Z digest=sha256:425213384842d1d201158c7b086d420f4d749abf2842272a235ae6d1b3a8ba24

Observation b4a91c25-d32e-4a6f-a5d3-5e61ae026181 · outbound

This paper cites @ crepe: Can vision-language foundation models reason compositionally?.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework @ crepe: Can vision-language foundation models reason compositionally?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.293183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.293183Z digest=sha256:9a5778f9d683affade9e3644de05695674341215511113e5dc16828f92d03f7b

Observation 4ebb616e-16a6-4190-a4bd-660078d895be · outbound

This paper cites Qwen2.5-vl,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Qwen2.5-vl,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.297452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.297452Z digest=sha256:28f935720f7729db948962818fa98f02a4ce8ad7581fe2055835d51f7713eeb9

Observation 1f87daaa-787a-4922-9028-3e10613dc958 · outbound

This paper cites 4o mini: Advancing cost-efficient intelligence, 2024,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework 4o mini: Advancing cost-efficient intelligence, 2024,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.838477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.301534Z digest=sha256:17c2ba8b090220d1c5cd07e458cb071edcd10006c68a298fa5e0c3baf341d4f0

Observation 39ae7dd0-f211-4e54-874e-c68fc8233f56 · outbound

This paper cites Hello gpt-4o,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Hello gpt-4o,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.823122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.305912Z digest=sha256:7a337bbb21c8a1fbb58928ac2f4bef0af3f344ff33b04c2b96548e47f25e11e0

Observation b03738ee-646c-4502-8f5f-188b317569f9 · outbound

This paper cites Introducing gpt-5,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Introducing gpt-5,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.807507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.310555Z digest=sha256:1f85a1bcc4bd0c4a6573c1a2e66cee891ee3f9e4201b16d8670bf25e18fc1291

Observation 51f45e83-d199-4222-bc23-0a51e053ed2b · outbound

This paper cites Learning to localize objects improves spatial reasoning in visual-llms,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Learning to localize objects improves spatial reasoning in visual-llms,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.791861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.314572Z digest=sha256:969353397a1128031d385a6f280bed654bbd442184efd793226caefc461c1a60

Observation 8a1837d5-e9a4-484d-a90b-0c32bec08abb · outbound

This paper cites HiMix: Reducing Computational Complexity in Large Vision-Language Models.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework HiMix: Reducing Computational Complexity in Large Vision-Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.497001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.318565Z digest=sha256:d1c99bf4d05d54790b7271edf7d2aa7a2fe549d2ba6678f32ccf1c084efe9b24

Observation d453f790-a4cd-49b5-b5db-c2bb756e6185 · outbound

This paper cites Eliminating the language bias for visual question answering with fine- grained causal intervention,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Eliminating the language bias for visual question answering with fine- grained causal intervention,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.774850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.323075Z digest=sha256:d7b18230375a1913a065d4a299eefb2d674e25ffdc5f1eb25fc70001646c81f1

Observation bbc7420b-076e-4705-9814-f9e1961ef556 · outbound

This paper cites Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.327797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.327797Z digest=sha256:ecabc7086e5cac986ebd4594dc7e10a3e51f9ef78230c13014bfcb5011059669

Observation e538e280-5cfb-46b7-a517-36b48a81d7a5 · outbound

This paper cites Diversify, Rationalize, and Combine: Ensembling Multiple QA Strategies for Zero-shot Knowledge-based VQA.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Diversify, Rationalize, and Combine: Ensembling Multiple QA Strategies for Zero-shot Knowledge-based VQA

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.459183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.333286Z digest=sha256:165c82e965e15918d96aae22e99df70afb4e7958bc1e3dabfe5267213bb61c5c

Observation cfe634d1-97ed-4d42-a49c-e3f291b3d187 · outbound

This paper cites Harnessing the Power of Multi-Task Pretraining for Ground-Truth Level Natural Language Explanations.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Harnessing the Power of Multi-Task Pretraining for Ground-Truth Level Natural Language Explanations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.338237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.338237Z digest=sha256:01178e04fb663ef01106451854b2e9cb34a496fb47ed56efbe39371a9b81ce12

Observation 614997c3-fa05-4221-81c6-761128f6e041 · outbound

This paper cites Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language Compositionality.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language Compositionality

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.436188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.343259Z digest=sha256:0cf9bf7ff11b1454d3c4f073ed701609c584db0943181b1081e06de01a0b27e0

Observation 886cfc9d-39ab-4256-aaa3-4f24b4803d38 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.348354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.348354Z digest=sha256:828f58d6d53251341e94663b2f599f3229bf0515082ea90c798f6c3f0ed4286f

Observation 0426853d-c62e-4887-8569-5ba81f7ec38d · outbound

This paper cites Available: https://openreview.net/forum?id=L4nOxziGf9.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Available: https://openreview.net/forum?id=L4nOxziGf9

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.168000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-09T18:03:53.152560Z digest=sha256:0ec65997e8ab94b07e00920aba4859106b9d66afd18f10c3e55997f9cf89cbde

Pith citing papers

Observation 798e4ab8-facb-4bdd-b162-b30819d9d17f · inbound

Augmented Vision-Language Models: A Systematic Review cites this paper.

Augmented Vision-Language Models: A Systematic Review VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework

Reference 129

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:46.154777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T14:33:44.595993Z digest=sha256:740bea4352b8aebd1a324a08be7263235d1a0c7d3ddf6770a88038911e5273c5