Pith. sign in

Paper Citation Record · LEDGER

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework

As of 10 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2502.00711.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00711 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:03:53.348354Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:33:44.595993Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T14:33:46.092736Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact6
  • verified fuzzy30
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 649cd177-51ed-4f04-9e08-d8538d083313 · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.108911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.108911Z digest=sha256:80a18a36d76e56c08b520ad6be951d75ba1e75789a74ed8954428a8775c91185

Observation d1e22241-9f19-4b57-a981-2e8df316fd40 · outbound

This paper cites Visual instruction tuning,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Visual instruction tuning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.114192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.114192Z digest=sha256:35342153f7cee101d69b340de06562e8b22da51470043cc300b8cb3f107b7027

Observation ddb8f5b2-24eb-467b-9b82-dc9abbc70039 · outbound

This paper cites Mdetr-modulated detection for end-to-end multi-modal understanding,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Mdetr-modulated detection for end-to-end multi-modal understanding,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.273883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.118804Z digest=sha256:deff0f9947db1e30bdf6326cd8c4466953ad15020256bfd700c5d31617daf568

Observation cad07e75-4fab-4b70-b696-a2df3173f6ad · outbound

This paper cites Omni-smola: Boosting generalist multimodal models with soft mixture of low-rank experts,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Omni-smola: Boosting generalist multimodal models with soft mixture of low-rank experts,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.259335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.123583Z digest=sha256:56826a7e2b3cec47f488b9e6c996f48bb20604f91f03bfae9de727fe207def5b

Observation 78d15b40-0858-4e91-b723-e66af2381bdf · outbound

This paper cites Cola: A benchmark for compositional text-to-image retrieval,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Cola: A benchmark for compositional text-to-image retrieval,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.244820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.128033Z digest=sha256:fedf8fc4f279fb47a4116bf8da133f6ab3f3c6e2f06c87fb6ef82f92f08e62ba

Observation fd748ad8-7833-406c-b731-c2e9de31d20c · outbound

This paper cites Visual programming: Compositional visual reasoning without training,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Visual programming: Compositional visual reasoning without training,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.230148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.132829Z digest=sha256:d7d0ad843618e2addf3e4a3ed9fa1da555ff8cf2418d00904c767011e8b40404

Observation 13ca5d95-21ba-4f0c-83a7-a6ef98b275c2 · outbound

This paper cites Interpretable visual reasoning: A survey,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Interpretable visual reasoning: A survey,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.216095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.138224Z digest=sha256:fb5a1ae7d24b9dc13305a0fe11c7b101f519dd02b31dc0ca15bf7897c130cc2d

Observation 9d897cae-9bd3-4bce-925a-d4d70a53e570 · outbound

This paper cites Rapper: Reinforced rationale-prompted paradigm for natural language explanation in visual question answering,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Rapper: Reinforced rationale-prompted paradigm for natural language explanation in visual question answering,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.200839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.143782Z digest=sha256:12640f36e03497e274577e9aa3bd0220c91d3d02b7088eccbcfcb2bc0ee7483a

Observation 87d224cb-6bc4-47bf-90d1-cea4ba5904f5 · outbound

This paper cites Rephrase, augment, reason: Visual grounding of questions for vision-language models,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Rephrase, augment, reason: Visual grounding of questions for vision-language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.183457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.148002Z digest=sha256:45fb383bc2cfbf35dd61a054d4e07527fa67a71ee8b7258515d30dcde191adfb

Observation 9c97c9b4-04b5-43ad-876c-31afd58a9186 · outbound

This paper cites Toward multi-granularity decision- making: Explicit visual reasoning with hierarchical knowledge,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Toward multi-granularity decision- making: Explicit visual reasoning with hierarchical knowledge,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.152331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.157639Z digest=sha256:169bada1ebd82b5a112c79205a0d85faa8c3f89ddeb4721bb1fc06f89a5e116b

Observation e6cc5687-cb05-4ac4-89df-ad9c12f108cb · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.162330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.162330Z digest=sha256:1944d2838d7e5824d98abeb4c6d52e3b25c79870cb47ac369e83767a5f0396cc

Observation 4ecd5353-0806-429f-bd7d-adc11a0573bb · outbound

This paper cites Vision–language model for visual question answering in medical imagery,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Vision–language model for visual question answering in medical imagery,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.137221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.167035Z digest=sha256:20fc42b56a660901ed4f6f53c3ae4202eec32098f682ae3e31a185c21b4b5b05

Observation 09bf68a5-9800-46fb-a189-f9a5aa0935ad · outbound

This paper cites Lingoqa: Visual question answering for autonomous driving,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Lingoqa: Visual question answering for autonomous driving,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.120736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.172078Z digest=sha256:b253cf6da9c4d55f6595e5ff79b2c86c94b3d412eb2db7b3246e9c62d844eb60

Observation 493c5b92-5199-4a6a-9220-f8dc3ab48917 · outbound

This paper cites VQA and Visual Reasoning: An Overview of Recent Datasets, Methods and Challenges.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework VQA and Visual Reasoning: An Overview of Recent Datasets, Methods and Challenges

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.176903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.176903Z digest=sha256:9af4fb2ffb175bb461ceaf29fe0672b8c11ff78c78fa7e7568aa1f753d404f03

Observation a8ac84f1-acd6-40fd-a085-f22403b49a64 · outbound

This paper cites Dealing with Semantic Underspecification in Multimodal NLP.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Dealing with Semantic Underspecification in Multimodal NLP

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.181943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.181943Z digest=sha256:51b81b2bb93a247f133dfdefc9fad4bc01ac693838909e49e03655e92ebb0b61

Observation 201d67f5-58d4-40b1-87a2-6cb46d878e3b · outbound

This paper cites Open visual knowledge extraction via relation-oriented multimodality model prompting,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Open visual knowledge extraction via relation-oriented multimodality model prompting,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.103734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.187050Z digest=sha256:368109d7e96ae9ccbdb8e06f99a3000751c51fcf277c6c22b8b93a95f06ae54e

Observation 45b14339-5b4d-41df-aaf5-dcf88d0e2447 · outbound

This paper cites PV2TEA: Patching Visual Modality to Textual-Established Information Extraction.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework PV2TEA: Patching Visual Modality to Textual-Established Information Extraction

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.675681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.196382Z digest=sha256:63a060a3f661ba8e33c44293832d30e3f9005406f56579ca094e0bfced8c63a6

Observation c41710f0-af35-4fab-8584-779222491bc2 · outbound

This paper cites Recurrent fusion network for image captioning,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Recurrent fusion network for image captioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.073290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.201585Z digest=sha256:31a6ec9fe7623715241cff08e8cbd965d085fbdd7472a7dfef832ffaaafa7acc

Observation 9cc4f3b7-a6f7-44f6-a865-61ff0b598e07 · outbound

This paper cites Boosting image captioning with attributes,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Boosting image captioning with attributes,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.057716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.206094Z digest=sha256:8fd7525af5251aeab861fb05bb532d07f04776abf0fa78d9d327cf6439fb4bd8

Observation ec9a6f63-0b01-4463-8b84-a56fea172bf6 · outbound

This paper cites Promptcap: Prompt-guided image captioning for vqa with gpt-3,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Promptcap: Prompt-guided image captioning for vqa with gpt-3,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.042647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.210754Z digest=sha256:b4a16d9497664db42d0712d20ead452785724259023fbbe5e13cfb1281a55465

Observation 2b3de275-8ce0-4c20-9023-989fc64a8130 · outbound

This paper cites Injecting semantic concepts into end-to-end image captioning,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Injecting semantic concepts into end-to-end image captioning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.026025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.215491Z digest=sha256:9bdcc193660ca32fca815d224f5c452c314c0632f6fc3a570e8cde3c3a48718b

Observation 44a583f8-05d8-490e-9b87-d344108e3e9b · outbound

This paper cites Visual Commonsense based Heterogeneous Graph Contrastive Learning.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Visual Commonsense based Heterogeneous Graph Contrastive Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.652545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.220310Z digest=sha256:88c9e37060c672f15a5621852a4b7cc05e191980b3735f903b0a7035f71d78fb

Observation 576ee7d7-bb2a-411f-bda0-bc7818ca20da · outbound

This paper cites Covlm: Composing visual entities and relationships in large language models via communicative decoding,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Covlm: Composing visual entities and relationships in large language models via communicative decoding,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.011138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.225287Z digest=sha256:5753d86d8906877adda80f7c208663ca980078acd2dcd761f4f4df28f3234266

Observation 859becd4-8893-49f5-8051-0a6a55c14891 · outbound

This paper cites Bridging knowledge graphs to generate scene graphs,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Bridging knowledge graphs to generate scene graphs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.995215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.229765Z digest=sha256:71fd6ed90f5adc55b2406262bc1a2a9868c76b14f91c2b16822fe22a4be36f0a

Observation 17da6dd9-5f85-4145-b893-db9cff8abcfb · outbound

This paper cites Co-training improves prompt-based learning for large language models,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Co-training improves prompt-based learning for large language models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.979908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.234699Z digest=sha256:acfbff1f23673b71921622c0d6a5cd5116c58b11df8a66228b39436b91d1153e

Observation 352cb6f8-0e86-4031-9d49-c0d03e1ee2c6 · outbound

This paper cites Large language model as attributed training data generator: A tale of diversity and bias,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Large language model as attributed training data generator: A tale of diversity and bias,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.964531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.239804Z digest=sha256:8a704074c7831dabc1a1f79636abaa98cbcbdaf56a480e3e4c93a1dbd09edc19

Observation f23dd74d-2d9e-4d04-965a-e696057e3844 · outbound

This paper cites Large language models are zero-shot reasoners,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Large language models are zero-shot reasoners,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.244544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.244544Z digest=sha256:d4241b62a903e800dc8aa7086cbc34ef93c0e62e8ad067e4d554682bb8561878

Observation 72faaa99-7ef1-4082-8c89-e4dd2dface47 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.249363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.249363Z digest=sha256:dbf51e645c361776373301ea899c0c9a0832d77196f953ad980bea8accabfa08

Observation 3bd099ee-9f1a-4352-b2f0-45ae8faa3e49 · outbound

This paper cites e-vil: A dataset and benchmark for natural language explanations in vision-language tasks,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework e-vil: A dataset and benchmark for natural language explanations in vision-language tasks,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.937640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.254268Z digest=sha256:5e98f2dda0b34aea1a7220417a6330d22281fdd8c2ae28d30f421067c47db6e2

Observation 37485df7-ccff-4907-a53b-d104118fabc5 · outbound

This paper cites NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.411763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.258591Z digest=sha256:5f38ee4aa31f9ce39514d5b816cca6e7f7ed89c031c8e203e2f5855c43f332e9

Observation 7e32c95e-1594-4df8-a12c-5606734b071c · outbound

This paper cites Improving vision-and-language reasoning via spatial relations modeling,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Improving vision-and-language reasoning via spatial relations modeling,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.088306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.262983Z digest=sha256:e118cd44f6f52ab1ea34f979a6945d7a5b1629c8bac86d78b1f4e889164702c6

Observation 581bbdad-363a-4dde-a6dd-c3153a860408 · outbound

This paper cites GPT-4 Technical Report.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework GPT-4 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.267186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.267186Z digest=sha256:bb82fe0e84aa46bf575d1d78716722405333adeaa9d0059ffd06c58f4cb9d169

Observation 49a49c73-976b-48ed-b03b-edf5ca7db438 · outbound

This paper cites Chain of thought prompting elicits reasoning in large language models,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Chain of thought prompting elicits reasoning in large language models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.920798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.271528Z digest=sha256:3131b04bbf8755190b61b31a2d8a3535ea4e87b0c01250be6a13be9df7303e93

Observation 6301d550-8086-41db-9af7-2e5291357021 · outbound

This paper cites Re- flexion: Language agents with verbal reinforcement learning,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Re- flexion: Language agents with verbal reinforcement learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.276014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.276014Z digest=sha256:dd0baf302ad00a369faeb3e2ee56167b15f095e506ccbbbf176fb0aa644a6d99

Observation 48cb50bd-3dce-4d0a-b28a-5375e5202022 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.280509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.280509Z digest=sha256:59030bbc1240833c20662a265e30353d8d0478a1aedfaec55c577381f2ff35b4

Observation ffc1b197-7c6f-4a7c-ab72-dbc64b4aa10b · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework A-okvqa: A benchmark for visual question answering using world knowledge,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.883686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.284831Z digest=sha256:92fbb4a03a58de1446f60d16f4ebec9482576a1189f44dff792bdf2781280245

Observation 2c7d8b91-5454-4649-8acd-ffe132606bac · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Vizwiz grand challenge: Answering visual questions from blind people,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.866801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.288845Z digest=sha256:94bb362ad6c7e2ba15eed6445da1991f69e7558f9045ca5bf4084081f740ddff

Observation b4a91c25-d32e-4a6f-a5d3-5e61ae026181 · outbound

This paper cites @ crepe: Can vision-language foundation models reason compositionally?.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework @ crepe: Can vision-language foundation models reason compositionally?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.293183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.293183Z digest=sha256:1b118281660d5bc4c96ca47118593aeb25a258547fce551fe9663de45d5dfda8

Observation 4ebb616e-16a6-4190-a4bd-660078d895be · outbound

This paper cites Qwen2.5-vl,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Qwen2.5-vl,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.297452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.297452Z digest=sha256:ec5fec6c5496aef20f9891be6fc258008eee1b08accf8d5cb0d8f673394947ca

Observation 1f87daaa-787a-4922-9028-3e10613dc958 · outbound

This paper cites 4o mini: Advancing cost-efficient intelligence, 2024,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework 4o mini: Advancing cost-efficient intelligence, 2024,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.838477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.301534Z digest=sha256:d316384dc0ec4cdad87cdeb88e75c599112ce90de7dcf07f98b1a39ca9aa6748

Observation 39ae7dd0-f211-4e54-874e-c68fc8233f56 · outbound

This paper cites Hello gpt-4o,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Hello gpt-4o,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.823122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.305912Z digest=sha256:3626c1f715745dc9df5dbb1ae71483f3b57338548429c6e3ccfeae0e05fe2e4f

Observation b03738ee-646c-4502-8f5f-188b317569f9 · outbound

This paper cites Introducing gpt-5,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Introducing gpt-5,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.807507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.310555Z digest=sha256:deee1334290bd4787fba732ad17daeac4890201c6c5b0090982b38d75485aa78

Observation 51f45e83-d199-4222-bc23-0a51e053ed2b · outbound

This paper cites Learning to localize objects improves spatial reasoning in visual-llms,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Learning to localize objects improves spatial reasoning in visual-llms,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.791861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.314572Z digest=sha256:0b822e15a3f55022661100465399d9ab392fcc552156b4fd7c311003e905ca3c

Observation 8a1837d5-e9a4-484d-a90b-0c32bec08abb · outbound

This paper cites HiMix: Reducing Computational Complexity in Large Vision-Language Models.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework HiMix: Reducing Computational Complexity in Large Vision-Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.497001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.318565Z digest=sha256:3c803c78bd4fb0a582cffac7f7dd00b28df874f70ba902545bd67d5e386bdb96

Observation d453f790-a4cd-49b5-b5db-c2bb756e6185 · outbound

This paper cites Eliminating the language bias for visual question answering with fine- grained causal intervention,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Eliminating the language bias for visual question answering with fine- grained causal intervention,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.774850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.323075Z digest=sha256:56dd972a0c20be46e5d34a4b39dc3ca246dceea00e3c6e1df9098953814c3de7

Observation bbc7420b-076e-4705-9814-f9e1961ef556 · outbound

This paper cites Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.327797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.327797Z digest=sha256:efda0c219f60ce1104fabd1abe61bbed8672e35cab6df1483c4da20df6185771

Observation e538e280-5cfb-46b7-a517-36b48a81d7a5 · outbound

This paper cites Diversify, Rationalize, and Combine: Ensembling Multiple QA Strategies for Zero-shot Knowledge-based VQA.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Diversify, Rationalize, and Combine: Ensembling Multiple QA Strategies for Zero-shot Knowledge-based VQA

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.459183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.333286Z digest=sha256:da286ec1add663f1581b6db973baca92b495746e4698185f48eb67490fc8448f

Observation cfe634d1-97ed-4d42-a49c-e3f291b3d187 · outbound

This paper cites Harnessing the Power of Multi-Task Pretraining for Ground-Truth Level Natural Language Explanations.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Harnessing the Power of Multi-Task Pretraining for Ground-Truth Level Natural Language Explanations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.338237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.338237Z digest=sha256:3c43e38e137fcb06d982da9529596dd1f2feecf11b130cd1a7f9b7122b08015d

Observation 614997c3-fa05-4221-81c6-761128f6e041 · outbound

This paper cites Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language Compositionality.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language Compositionality

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.436188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.343259Z digest=sha256:8e24b57ced2afccaccab1a6261a79d84d1c590203723057388f406e6fa7a075c

Observation 886cfc9d-39ab-4256-aaa3-4f24b4803d38 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.348354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.348354Z digest=sha256:73e3d72954a2c82706965bd3dac9cc219f81379e9a8a4a0d4e214d16db0e1b85

Observation 0426853d-c62e-4887-8569-5ba81f7ec38d · outbound

This paper cites Available: https://openreview.net/forum?id=L4nOxziGf9.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Available: https://openreview.net/forum?id=L4nOxziGf9

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.168000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.152560Z digest=sha256:10dc41dcee6b66608b746119c4fe4c9519d9c6dab9f51b94d512b62611888b73

Pith citing papers

Observation 798e4ab8-facb-4bdd-b162-b30819d9d17f · inbound

Augmented Vision-Language Models: A Systematic Review cites this paper.

Augmented Vision-Language Models: A Systematic Review VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework

Reference 129

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:46.154777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:33:44.595993Z digest=sha256:711598f637394913af0a11a9fa9ec9eba479a320322c9f8aa30d33d49e460396