Pith. sign in

Paper Citation Record · LEDGER

Augmented Vision-Language Models: A Systematic Review

As of 8 August 2026, this Paper Citation Record lists 100 of 135 outbound references and 0 inbound Pith citation observations for arXiv:2507.22933.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22933 v1

Coverage vector

measured 100 of 135 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:33:42.097727Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 135 outbound references displayed

  • verified exact21
  • verified fuzzy0
  • unresolved74
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b5ed89e-d015-4e59-bdd0-47f05ff22454 · outbound

This paper cites Spatial Knowledge Distillation to aid Visual Reasoning.

Augmented Vision-Language Models: A Systematic Review Spatial Knowledge Distillation to aid Visual Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:29.127796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:29.127796Z digest=sha256:5a55537cd2e0f3b3b0663c6f241be088272f8169a1b97d5e58e0fc5d712c9f24

Observation b7516853-0591-4f1b-9709-9794466a7b12 · outbound

This paper cites AQuA: ASP-Based Visual Question Answering.

Augmented Vision-Language Models: A Systematic Review AQuA: ASP-Based Visual Question Answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:29.645049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:29.645049Z digest=sha256:9a30e9459b6b00a60afc66d16a1da122a9bf6e52c1a88a308395854e69425ef9

Observation 8b231a0e-fdb1-4419-b9fa-67498c27854b · outbound

This paper cites Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models.

Augmented Vision-Language Models: A Systematic Review Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:30.145531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:30.145531Z digest=sha256:85ad726ed42f54ccaeef489087de573cb2404cf579b1bddb66975025c2da3e2f

Observation cde486f6-7cac-4a0d-b1be-2b8804c12853 · outbound

This paper cites Thinking fast and slow in ai.

Augmented Vision-Language Models: A Systematic Review Thinking fast and slow in ai

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:30.284595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:30.284595Z digest=sha256:1dd4c0d9356fc97a0099fd6205051670cb0731adc14e81ef7d7855f1d774eb12

Observation 35c08563-9707-4a6a-bf06-4c6b215050f0 · outbound

This paper cites Linguistically routing capsule network for out-of-distribution visual question answering.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp.

Augmented Vision-Language Models: A Systematic Review Linguistically routing capsule network for out-of-distribution visual question answering.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:30.582183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:30.582183Z digest=sha256:a656c58cef768d53234385a1b1808eb461062828795e77247647f5de605681a5

Observation 35bc9da5-ae4e-488e-b46b-5a5207bba8d9 · outbound

This paper cites HAMMR: HierArchical MultiModal React agents for generic VQA.

Augmented Vision-Language Models: A Systematic Review HAMMR: HierArchical MultiModal React agents for generic VQA

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:30.949867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:30.949867Z digest=sha256:ef69b71e85fb51a82392f875212c19b91db19492080c33290351c0b3d7f099d7

Observation dc6cb8fe-6699-4bdb-81a8-c81c2b1476a9 · outbound

This paper cites Multimodal Structured Generation: CVPR's 2nd MMFM Challenge Technical Report.

Augmented Vision-Language Models: A Systematic Review Multimodal Structured Generation: CVPR's 2nd MMFM Challenge Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:31.064740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:31.064740Z digest=sha256:6a499083c5687b2bf95efa452c9abf49f597ae4baee13b2851882de27f27d37f

Observation f4e1ee3b-f819-453c-93c6-5f8ffd097bd6 · outbound

This paper cites Retrieval Augmented Structured Generation: Business Document Information Extraction As Tool Use.

Augmented Vision-Language Models: A Systematic Review Retrieval Augmented Structured Generation: Business Document Information Extraction As Tool Use

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:31.242980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:31.242980Z digest=sha256:7ca848f75edb08d1f3c4564a3c5f126216404fa76472b56786a365fe49cd2166

Observation be8cbd0f-0794-4780-b72c-10e02c838b15 · outbound

This paper cites Uncertainty-based visual question answering: Estimating semantic incon- sistency between image and knowledge base.

Augmented Vision-Language Models: A Systematic Review Uncertainty-based visual question answering: Estimating semantic incon- sistency between image and knowledge base

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:31.445980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:31.445980Z digest=sha256:67f9d057b3ff8e27accd439202c30a013b2cb7a166b059224847eb9bdda879ef

Observation a7d9639e-f645-4088-80e7-25eeee52864d · outbound

This paper cites Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?.

Augmented Vision-Language Models: A Systematic Review Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:31.536273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:31.536273Z digest=sha256:aecaeb5e6d83a5decf912bdefae600bfb74bcafa19d267ef4d100a7ba920ac14

Observation 5a94369f-d9e9-4c41-8409-346b642dc0e8 · outbound

This paper cites Zero-shot Visual Question Answering using Knowledge Graph.

Augmented Vision-Language Models: A Systematic Review Zero-shot Visual Question Answering using Knowledge Graph

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:31.629947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:31.629947Z digest=sha256:ee61f86ddc1d07f20d2fe72a3c4686448b3054d3e96b751d7b50c2c6052fb414

Observation 81f4ea05-5cd9-4796-9a59-6af62f1840fe · outbound

This paper cites MORE: Multi-mOdal REtrieval Augmented Generative Commonsense Reasoning.

Augmented Vision-Language Models: A Systematic Review MORE: Multi-mOdal REtrieval Augmented Generative Commonsense Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:31.755675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:31.755675Z digest=sha256:a29eff56561a54e938370fed86b45e20b3d4eb6c65ee90fda83771559cd63bb7

Observation a7405663-718a-43c1-89a3-d915c0f44617 · outbound

This paper cites The Role of Foundation Models in Neuro-Symbolic Learning and Reasoning.

Augmented Vision-Language Models: A Systematic Review The Role of Foundation Models in Neuro-Symbolic Learning and Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:31.919589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:31.919589Z digest=sha256:22459cec94e15ca28f179373776116ad835f3baeca1831e5bcd19616f6991f4a

Observation 10bb64b2-80d6-411f-bbb3-58e3c1caa9b5 · outbound

This paper cites EKTVQA: Generalized use of External Knowledge to empower Scene Text in Text-VQA.

Augmented Vision-Language Models: A Systematic Review EKTVQA: Generalized use of External Knowledge to empower Scene Text in Text-VQA

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:32.034495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:32.034495Z digest=sha256:9edd37bda4683aeede6c60ba1c4b47c7179e285fcec86c6beafb983a302a75db

Observation 75d6735e-c162-4155-b849-1ab0ee2781fb · outbound

This paper cites Multimodal Foundation Models for Zero-shot Animal Species Recognition in Camera Trap Images.

Augmented Vision-Language Models: A Systematic Review Multimodal Foundation Models for Zero-shot Animal Species Recognition in Camera Trap Images

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:32.099441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:32.099441Z digest=sha256:948a9d3db92790f54acf060849d84aa423e56353871c6e4d8c60b9eb5215ef31

Observation 13f78976-f146-4f2e-acd1-a6bf30837265 · outbound

This paper cites VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding.

Augmented Vision-Language Models: A Systematic Review VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:32.198959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:32.198959Z digest=sha256:9447100a72e369ef866e084a32092f28511d8903831628040edd8b79b41320c7

Observation 7e197d21-f03c-46e4-9226-cf6f7983a077 · outbound

This paper cites Generate then Select: Open-ended Visual Question Answering Guided by World Knowledge.

Augmented Vision-Language Models: A Systematic Review Generate then Select: Open-ended Visual Question Answering Guided by World Knowledge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:32.228324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:32.228324Z digest=sha256:8cc278eb42ac96394b5a0f8496571df6e9a28b4ec7bd244a8eb5e0f03227e9c6

Observation c4d81b7f-a88f-4521-8290-c6bd12809865 · outbound

This paper cites Open-Set Knowledge-Based Visual Question Answering with Inference Paths.

Augmented Vision-Language Models: A Systematic Review Open-Set Knowledge-Based Visual Question Answering with Inference Paths

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:32.365450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:32.365450Z digest=sha256:f3e475f7fcf91e4d7a63a1cac821b7908bbd32f0088d37fc3e72525d58a2804a

Observation 2b75d489-02a2-45ee-9225-1c2042dd47bf · outbound

This paper cites CRIC: A VQA Dataset for Compositional Reasoning on Vision and Commonsense.

Augmented Vision-Language Models: A Systematic Review CRIC: A VQA Dataset for Compositional Reasoning on Vision and Commonsense

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:32.517392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:32.517392Z digest=sha256:b223a7b9d1e80be071ecf54a5171c159c1cf584d837e6e5de60ed3e4b5fe6434

Observation b47f70d0-0c4d-448d-a1ae-c1e91af6e77f · outbound

This paper cites an unresolved cited work.

Augmented Vision-Language Models: A Systematic Review Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:32.640882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:32.640882Z digest=sha256:d1555749bdf7f20cc0b94e1f94d41700fe51064e62d07f151fd231748e957ade

Observation d7da0995-7e82-4702-9077-8d170970519a · outbound

This paper cites Cross-modal object detection based on a knowledge update.Sensors (Basel, Switzerland), 22, 2022b.

Augmented Vision-Language Models: A Systematic Review Cross-modal object detection based on a knowledge update.Sensors (Basel, Switzerland), 22, 2022b

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:32.798914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:32.798914Z digest=sha256:ecbe523453e16c7ba2c18591b089b7ca8e1b112929aedcc6b82a305a8f1248e9

Observation 242ef5ae-5af6-455d-bc93-4c5823daedab · outbound

This paper cites Scene Graph Generation with External Knowledge and Image Reconstruction.

Augmented Vision-Language Models: A Systematic Review Scene Graph Generation with External Knowledge and Image Reconstruction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:32.838823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:32.838823Z digest=sha256:9b0d315c17aff11c96cd81061d7b071d95e5a7a5a448dc3623dca37a798a190c

Observation c52725fd-73e1-406e-99ad-cbe082c6f92f · outbound

This paper cites KAT: A Knowledge Augmented Transformer for Vision-and-Language.

Augmented Vision-Language Models: A Systematic Review KAT: A Knowledge Augmented Transformer for Vision-and-Language

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:32.886456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:32.886456Z digest=sha256:e60de7f4896b663f7bbc917cc6f0c5c0b38f38e9b7ab261fe69e1d3501af3115

Observation d3a5bc0b-4cf7-4701-9138-2d6a1f413096 · outbound

This paper cites Visual Programming: Compositional visual reasoning without training.

Augmented Vision-Language Models: A Systematic Review Visual Programming: Compositional visual reasoning without training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:32.978926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:32.978926Z digest=sha256:b4ba13e6d27636212579580f0dd4697697529ebe300cf486172089fcc8ef3bbe

Observation 305b795c-83da-47fc-9c41-ac2cd71274d4 · outbound

This paper cites Cross-Modal Retrieval Augmentation for Multi-Modal Classification.

Augmented Vision-Language Models: A Systematic Review Cross-Modal Retrieval Augmentation for Multi-Modal Classification

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:33.090097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:33.090097Z digest=sha256:3d8751f65584ea319fa200c1c690e92d6bcc900adc06350f4535a195dabc8547

Observation af157ac8-0280-4633-8fb1-b71a1496c831 · outbound

This paper cites Knowledge Condensation and Reasoning for Knowledge-based VQA.

Augmented Vision-Language Models: A Systematic Review Knowledge Condensation and Reasoning for Knowledge-based VQA

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:33.190135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:33.190135Z digest=sha256:aec7a05ad0018af3c5d9375d71e0159715be2c6fee1a50d4eb100d57bd4534c9

Observation 5f7f103d-d20e-404a-8951-0119b546812a · outbound

This paper cites Hypergraph Transformer: Weakly-supervised Multi-hop Reasoning for Knowledge-based Visual Question Answering.

Augmented Vision-Language Models: A Systematic Review Hypergraph Transformer: Weakly-supervised Multi-hop Reasoning for Knowledge-based Visual Question Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:33.309138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:33.309138Z digest=sha256:1f75a474fe8856189a69febcc5213c4ccb26693ecbf1420234436decc410fba3

Observation 3c12baba-5d6d-4208-bbee-9f3a37ea0b62 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Augmented Vision-Language Models: A Systematic Review Training Compute-Optimal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:33.424386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:33.424386Z digest=sha256:6b356e15ab6a7f1eb14a4c2ee107391d30e3369574edf67b0ac8f38c96e1e029

Observation fcee05e0-2bca-4e21-a978-d0b43612fd93 · outbound

This paper cites Tool Documentation Enables Zero-Shot Tool-Usage with Large Language Models.

Augmented Vision-Language Models: A Systematic Review Tool Documentation Enables Zero-Shot Tool-Usage with Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:33.565983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:33.565983Z digest=sha256:3b6e49e6c513639aeec086ef0171bb71ac8d327a4f7c8dad09293e5872f033eb

Observation 474ce44a-9320-46cf-baa5-8e0de07e3743 · outbound

This paper cites Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

Augmented Vision-Language Models: A Systematic Review Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:33.739981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:33.739981Z digest=sha256:dda08379b3d0ba3b04d3dfff43f2ce2d729a5d503bcaa9105c14bc9f2e05c548

Observation 8b445b9f-da78-40c0-a183-7914780761c7 · outbound

This paper cites Schmid, David A.

Augmented Vision-Language Models: A Systematic Review Schmid, David A

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:33.845656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:33.845656Z digest=sha256:6f01440cb375b12e3fc9111c12e043cddff995bc3a31aa46618567872c0f073b

Observation 231d22e9-998e-416a-b493-c25990000920 · outbound

This paper cites AVIS: Autonomous Visual Information Seeking with Large Language Model Agent.

Augmented Vision-Language Models: A Systematic Review AVIS: Autonomous Visual Information Seeking with Large Language Model Agent

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:51.721138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:33.975118Z digest=sha256:bef7f5200480c54560923244c8e8dcc9cfc7d137c399d58cdab6ff163c4d04ac

Observation 60510c9d-9425-4d1b-b40b-6fdb131b6efe · outbound

This paper cites Learning by Abstraction: The Neural State Machine.

Augmented Vision-Language Models: A Systematic Review Learning by Abstraction: The Neural State Machine

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:34.144898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:34.144898Z digest=sha256:f5c8e8fb9b1db803eb86a03fea5cc7456bdf4ad8898ea1f7ad4ba1301281f358

Observation 50696577-8904-4c7e-a6df-2218803de19e · outbound

This paper cites Shahzad, and M.

Augmented Vision-Language Models: A Systematic Review Shahzad, and M

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:34.272947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:34.272947Z digest=sha256:8e1a5bd56f98080dbb1fcfb9009ebcbfc0ac486e4a387caf7df53a928dbb83eb

Observation b05febb3-bd38-4659-bc3d-102f49733d5c · outbound

This paper cites Retrieval-Enhanced Contrastive Vision-Text Models.

Augmented Vision-Language Models: A Systematic Review Retrieval-Enhanced Contrastive Vision-Text Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:51.551631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:34.426589Z digest=sha256:cbfb331aa1368242af42e860b2fed34553c49cd8ad0410809fa8df8649d1ce25

Observation 2660c230-13c2-4114-bfd0-4b69abccb51d · outbound

This paper cites Jamshed and M.

Augmented Vision-Language Models: A Systematic Review Jamshed and M

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:34.543307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:34.543307Z digest=sha256:f34597441ea9f4ba01486d80a2f40e1435931c70445a281ca8ae33a4e0632fd5

Observation db10c290-628d-4315-8122-b6f8750cf1bd · outbound

This paper cites Precision Empowers, Excess Distracts: Visual Question Answering With Dynamically Infused Knowledge In Language Models.

Augmented Vision-Language Models: A Systematic Review Precision Empowers, Excess Distracts: Visual Question Answering With Dynamically Infused Knowledge In Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:51.432525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:34.643711Z digest=sha256:8502b105b6a7a6bf589bf5442e32c91b2d6b6ffea1b940db29b9bf73e56b06f3

Observation 17076408-d210-4df5-abd3-7e1b2935b392 · outbound

This paper cites KAFA: Rethinking Image Ad Understanding with Knowledge-Augmented Feature Adaptation of Vision-Language Models.

Augmented Vision-Language Models: A Systematic Review KAFA: Rethinking Image Ad Understanding with Knowledge-Augmented Feature Adaptation of Vision-Language Models

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:33:51.256296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:34.754873Z digest=sha256:3cafcfb2b9dff6f0f4fa30243cf332a3b21e5e8538226c5cf3f8d4db120bdc5a

Observation 355fc88c-1278-4f15-a50b-425e4a6823f1 · outbound

This paper cites an unresolved cited work.

Augmented Vision-Language Models: A Systematic Review Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:34.883568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:34.883568Z digest=sha256:f953e4554445be3940dd4113c4420de559de7ba40591b26b68318e98bafe9d07

Observation 84226f4a-59d7-4408-9bb8-c55828f11079 · outbound

This paper cites How Can We Know What Language Models Know?.

Augmented Vision-Language Models: A Systematic Review How Can We Know What Language Models Know?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:34.997393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:34.997393Z digest=sha256:d6c39c641b4f7cb747bc24f721652b6f51d9b5f5f10a0751a1de19ec4c1d0067

Observation 67d66ebf-a1ef-4fc6-93d4-edb07ff9c1a8 · outbound

This paper cites CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning.

Augmented Vision-Language Models: A Systematic Review CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:35.054365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:35.054365Z digest=sha256:c0c1df3cc758d72e859d525f1c9e9499bc5450a29877a40da9334655831b8969

Observation 94452ba9-1d03-4dad-a326-b38f60a5fa29 · outbound

This paper cites Knowledge-aware prompt tuning for generalizable vision-language models.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pp.

Augmented Vision-Language Models: A Systematic Review Knowledge-aware prompt tuning for generalizable vision-language models.2023 IEEE/CVF International Conference on Computer Vision (ICCV), pp

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:35.176470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:35.176470Z digest=sha256:bde07c5699c5a3ee2095b1b07c9e886cb01ef422dddfb517557aa496e1bbc5f9

Observation 7a11f80c-2da5-4677-9d7e-9fddc4b4fcfd · outbound

This paper cites Scaling Laws for Neural Language Models.

Augmented Vision-Language Models: A Systematic Review Scaling Laws for Neural Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:35.301992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:35.301992Z digest=sha256:6ede35372588518c2b38eb50af9dfc483f2c8894bc83f5f3ba9f7239659277ea

Observation f99d6b87-f9a2-4066-8d17-5af24235b8b4 · outbound

This paper cites RAGAR, Your Falsehood Radar: RAG-Augmented Reasoning for Political Fact-Checking using Multimodal Large Language Models.

Augmented Vision-Language Models: A Systematic Review RAGAR, Your Falsehood Radar: RAG-Augmented Reasoning for Political Fact-Checking using Multimodal Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:35.436289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:35.436289Z digest=sha256:747806ee1346fbf249fcc4392a5b1673f6ae480b32aad4c1679cdea4c8d86012

Observation 902e3389-e2ad-4cfd-9a77-25037ab00231 · outbound

This paper cites Analyzing Modular Approaches for Visual Question Decomposition.

Augmented Vision-Language Models: A Systematic Review Analyzing Modular Approaches for Visual Question Decomposition

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:35.576006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:35.576006Z digest=sha256:fe78b11779634dca1e8536d98cf66735f10e9e5a3f2b1f4622a75c0edd784afd

Observation f0e62834-e335-45c8-a7e1-4bf4b6d4e485 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

Augmented Vision-Language Models: A Systematic Review VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:35.672466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:35.672466Z digest=sha256:316114820db59acdf5f41107396b7288ff9eb8802c1fabc5e5ba082c936b88dc

Observation 30899bee-993b-46e9-9f36-8f52f16528e6 · outbound

This paper cites Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations.

Augmented Vision-Language Models: A Systematic Review Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:35.781511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:35.781511Z digest=sha256:d61044ec65783d2ca1bea252d63587491a14ebe7c4a232caccbc336b008f9fcb

Observation ee10abd9-2a54-4e3a-a0bb-c57f26c97a14 · outbound

This paper cites Introduction to Soar.

Augmented Vision-Language Models: A Systematic Review Introduction to Soar

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:35.923364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:35.923364Z digest=sha256:6756a03ff84d7310be818cf3adf192c50a2288f80e3d909c57d65eb88225be64

Observation 428497d6-5ca8-41f1-8b8d-9e302d79dd67 · outbound

This paper cites Multimodal Reasoning with Multimodal Knowledge Graph.

Augmented Vision-Language Models: A Systematic Review Multimodal Reasoning with Multimodal Knowledge Graph

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:36.019772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:36.019772Z digest=sha256:95a880300cf8e3ca68dd236af5cd5803e8e33b305348552b532db24eae4f0c07

Observation 141cd059-487c-4696-8fa5-f7180f3eef0c · outbound

This paper cites Incorporating External Knowledge to Answer Open-Domain Visual Questions with Dynamic Memory Networks.

Augmented Vision-Language Models: A Systematic Review Incorporating External Knowledge to Answer Open-Domain Visual Questions with Dynamic Memory Networks

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:50.923535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:36.210167Z digest=sha256:24342af16221b982627244d896ab61e624290f93c283c51208c0996c9f585431

Observation 67518e60-73e7-464c-bafd-eef55c592d11 · outbound

This paper cites Visual Question Answering as Reading Comprehension.

Augmented Vision-Language Models: A Systematic Review Visual Question Answering as Reading Comprehension

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:50.733235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:36.369985Z digest=sha256:dcd674d612a219b883d055d1cb547aef9184889925679f7eb38123bceb682d04

Observation e0dbbc37-0153-40f7-a113-993d10f73528 · outbound

This paper cites Supporting Vision-Language Model Inference with Confounder-pruning Knowledge Prompt.

Augmented Vision-Language Models: A Systematic Review Supporting Vision-Language Model Inference with Confounder-pruning Knowledge Prompt

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:50.584470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:36.489057Z digest=sha256:8c5dc5ce6a10be8937b849316f7186fb0a89bd08ae49a391564e1eef828f913b

Observation 13061f17-06f0-4bd2-9439-ad74f130ac88 · outbound

This paper cites GraphAdapter: Tuning Vision-Language Models With Dual Knowledge Graph.

Augmented Vision-Language Models: A Systematic Review GraphAdapter: Tuning Vision-Language Models With Dual Knowledge Graph

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:33:50.470486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:36.593062Z digest=sha256:c79fb4ad5f58b74ab3ae5ffd324242a19167655198fa78ad2fd94217f52231d9

Observation 6dca3c25-287b-40ed-9bc5-f95a92bdbfcb · outbound

This paper cites Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning.

Augmented Vision-Language Models: A Systematic Review Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:50.306626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:36.708293Z digest=sha256:3f4799305a03ef1f660c8a0bd3463a8e978d654992aa51f09f187d3d1e3f8ae2

Observation 9dc30605-c63c-42ec-8e71-f099eebd1f2a · outbound

This paper cites Maria: A Visual Experience Powered Conversational Agent.

Augmented Vision-Language Models: A Systematic Review Maria: A Visual Experience Powered Conversational Agent

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:50.142739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:36.793809Z digest=sha256:0b33d580d07edbbb90a131e4a5290723019f828b6dff0bc4d4725586aba83528

Observation 0008139f-312f-44ea-85ae-b4c5eb729170 · outbound

This paper cites Towards Medical Artificial General Intelligence via Knowledge-Enhanced Multimodal Pretraining.

Augmented Vision-Language Models: A Systematic Review Towards Medical Artificial General Intelligence via Knowledge-Enhanced Multimodal Pretraining

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:33:50.007188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:36.879907Z digest=sha256:683730d2c87baf467da708ce9a232045828dfc6aa3d5d36c2bc3cc5370b79a84

Observation 3827c1f7-8752-4305-82b5-2721a1fcc27a · outbound

This paper cites Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question Answering.

Augmented Vision-Language Models: A Systematic Review Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question Answering

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:49.905173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:37.003011Z digest=sha256:2083bbc3b5230ddab7f4d10745a0898a5692f317bd42ee355f5c95e633d1051a

Observation 8de2450d-5709-4b81-83fe-8a020939aea5 · outbound

This paper cites Knowledge-Enhanced Hierarchical Information Correlation Learning for Multi-Modal Rumor Detection.

Augmented Vision-Language Models: A Systematic Review Knowledge-Enhanced Hierarchical Information Correlation Learning for Multi-Modal Rumor Detection

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:49.767228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:37.129951Z digest=sha256:3286849309391483895a9f5bced652b05046de85702cd63dc00c94167ef7170c

Observation 13d2176b-db56-4ac4-856e-40753e4f798d · outbound

This paper cites LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents.

Augmented Vision-Language Models: A Systematic Review LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:37.251420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:37.251420Z digest=sha256:2c4a5504139d93595fc85e2c7bcb8891d2be7b48c540c230fc6ee46e76d7453d

Observation c5512e85-dffb-4880-b213-28477b0d6502 · outbound

This paper cites Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models.

Augmented Vision-Language Models: A Systematic Review Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:37.359408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:37.359408Z digest=sha256:484247d8cd97f89d1d9cf833ebc5a92a187e39e404fa40c8d6ef32b112428a05

Observation 7045fb5e-9452-481a-a601-42663d5d60e6 · outbound

This paper cites The Next Decade in AI: Four Steps Towards Robust Artificial Intelligence.

Augmented Vision-Language Models: A Systematic Review The Next Decade in AI: Four Steps Towards Robust Artificial Intelligence

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:37.432857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:37.432857Z digest=sha256:5dff569f910c5dc0d830f486c31c8fafb7b9d0aa705a5bb61d2bea064f2b79ff

Observation ff7aa1ac-675a-4435-a58b-54fea219b2cf · outbound

This paper cites Gupta, and Marcus Rohrbach.

Augmented Vision-Language Models: A Systematic Review Gupta, and Marcus Rohrbach

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:37.550363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:37.550363Z digest=sha256:76ce236469bd4728f50a860c454e7b0c5e1d8769f0c3d732bcd9e02eb78a4fdc

Observation 08797055-7001-4854-8740-1abf11d14df3 · outbound

This paper cites Uijlings, Lluís Castrejón, A.

Augmented Vision-Language Models: A Systematic Review Uijlings, Lluís Castrejón, A

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:37.641589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:37.641589Z digest=sha256:109420608438a32d76d3d0dc764857b26d465b75672c93fa78acb19612557454

Observation 7b16d1ac-993a-43f0-a05e-64ff63a8c3bc · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Augmented Vision-Language Models: A Systematic Review GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:37.755569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:37.755569Z digest=sha256:225487ad13ba89b0120bb786fb36278597d5e94b82324ac8b5a22bd81aecd584

Observation 740733e0-1d8c-40ca-8e5d-26c0c27ea79a · outbound

This paper cites an unresolved cited work.

Augmented Vision-Language Models: A Systematic Review Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:37.878901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:37.878901Z digest=sha256:36bba2d7acb51b7816401a29d30c87ad75919ff2628bc096fd95e0029b6e6203

Observation c71edbd1-d4e4-4265-b0db-2f3ef20b880c · outbound

This paper cites VQA Training Sets are Self-play Environments for Generating Few-shot Pools.

Augmented Vision-Language Models: A Systematic Review VQA Training Sets are Self-play Environments for Generating Few-shot Pools

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:49.530589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:37.982939Z digest=sha256:08030b23fbd46529a95bd0de255830b761bde36b5b911e8dd01fa2b2cd0a0de2

Observation d7941ab1-49a2-4770-b62f-7faffac0ebb9 · outbound

This paper cites Fast Model Editing at Scale.

Augmented Vision-Language Models: A Systematic Review Fast Model Editing at Scale

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:38.133495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:38.133495Z digest=sha256:35191dcc2b0bba9898752bbe39ab1233041028e4b1ee90e413579d010cdd64e7

Observation 3390a54d-51f7-437b-b1ba-8ffedf266925 · outbound

This paper cites Hyper-dimensional computing for a visual question-answering system that is trainable end-to-end.

Augmented Vision-Language Models: A Systematic Review Hyper-dimensional computing for a visual question-answering system that is trainable end-to-end

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:49.412597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:38.220218Z digest=sha256:a32d9cffd73c6d84a1d76b2e063051874449b6ba6af1c71684030afb013db263

Observation 266efa4a-04cd-455f-9cb2-43637f5f826f · outbound

This paper cites Straight to the Facts: Learning Knowledge Base Retrieval for Factual Visual Question Answering.

Augmented Vision-Language Models: A Systematic Review Straight to the Facts: Learning Knowledge Base Retrieval for Factual Visual Question Answering

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:49.254650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:38.308600Z digest=sha256:f77e75e680891e893d7b0042fa5d5fdba0182e436b3895f36935cb5e0b8581b2

Observation 0373640a-fd1a-40b4-a802-de015ad8cef0 · outbound

This paper cites Out of the Box: Reasoning with Graph Convolution Nets for Factual Visual Question Answering.

Augmented Vision-Language Models: A Systematic Review Out of the Box: Reasoning with Graph Convolution Nets for Factual Visual Question Answering

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:49.102227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:38.465430Z digest=sha256:7e5d2eaede5c7803126b416a01d0ae1ff283c9d7b83fd1a9d0e27ab6cd3ac9ff

Observation 6b9c9dee-1b6b-438f-8cbf-93c910a3d696 · outbound

This paper cites External commonsense knowledge as a modality for social intelligence question-answering.

Augmented Vision-Language Models: A Systematic Review External commonsense knowledge as a modality for social intelligence question-answering

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:38.574871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:38.574871Z digest=sha256:449b4d4abc7752eeac3e09110484d5ae25e5aa792eca6105c3beace230a498f6

Observation 72882c68-7b0a-4fce-87b1-35b414ee6a69 · outbound

This paper cites ScreenAgent: A Vision Language Model-driven Computer Control Agent.

Augmented Vision-Language Models: A Systematic Review ScreenAgent: A Vision Language Model-driven Computer Control Agent

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:38.722140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:38.722140Z digest=sha256:5c6f102ed04af998cc1febeb796c5e94e4d2f91e1c2da13141ab5355ec56ff1d

Observation b9d173f6-613a-4894-bc9e-c2c4620f4a13 · outbound

This paper cites Prediction of actions and places by the time series recognition from images with multimodal llm.2024 IEEE 18th International Conference on Semantic Computing (ICSC), pp.

Augmented Vision-Language Models: A Systematic Review Prediction of actions and places by the time series recognition from images with multimodal llm.2024 IEEE 18th International Conference on Semantic Computing (ICSC), pp

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:38.816381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:38.816381Z digest=sha256:129516f4e21b61fed745968d4cd0b5bc5e9592f5c3293640d5a386d72c32089f

Observation d1a614cf-10cc-435a-8962-91b2d7a1c31f · outbound

This paper cites Modal-adaptive Knowledge-enhanced Graph-based Financial Prediction from Monetary Policy Conference Calls with LLM.

Augmented Vision-Language Models: A Systematic Review Modal-adaptive Knowledge-enhanced Graph-based Financial Prediction from Monetary Policy Conference Calls with LLM

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:48.975837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:38.895950Z digest=sha256:a7a1ae1ee782bfea3f89d75fbb692de76d071f8a943e8caf9841c95ad031ef29

Observation b5174dd7-3c72-417f-923b-8f7ed5e2c6fd · outbound

This paper cites Enhancing Cross-Modal Contextual Congruence for Crowdfunding Success using Knowledge-infused Learning.

Augmented Vision-Language Models: A Systematic Review Enhancing Cross-Modal Contextual Congruence for Crowdfunding Success using Knowledge-infused Learning

Reference 84

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:33:48.867608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:38.957536Z digest=sha256:71ef5cd5ceb7f853655adc4f49cff353a53f07c3a6354b69d41226d452d256db

Observation ad972237-8cf0-43c6-95d8-cc00e671cedd · outbound

This paper cites Generative Agents: Interactive Simulacra of Human Behavior.

Augmented Vision-Language Models: A Systematic Review Generative Agents: Interactive Simulacra of Human Behavior

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:39.057382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:39.057382Z digest=sha256:4de635cb695ff9bd0294c89dd80e321e011e0f82b54c313d164ac80da68466ae

Observation f0190988-7b3a-4d4c-803a-5a09aea532be · outbound

This paper cites Gorilla: Large Language Model Connected with Massive APIs.

Augmented Vision-Language Models: A Systematic Review Gorilla: Large Language Model Connected with Massive APIs

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:39.206412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:39.206412Z digest=sha256:0fb7e34ddb883782f89cc8681e6b75c3364a66d78404d4fa561fffbc773a9c57

Observation 029fae67-72e3-48c6-a9a4-7ee923bc1138 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

Augmented Vision-Language Models: A Systematic Review ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:39.386927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:39.386927Z digest=sha256:3c4466a8f2c916110a8edff76aa53cc79699416466765be9d8d96970edc6d1a0

Observation bb157cb8-2692-431f-8a0a-67c212cb6256 · outbound

This paper cites Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation.

Augmented Vision-Language Models: A Systematic Review Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:39.558793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:39.558793Z digest=sha256:45c421ee73f3b83e007ba6832636dc10735c163af55f894c08e359fed5616f50

Observation 7e9876d1-cd47-4057-b6e7-04b2e149e1b2 · outbound

This paper cites Ksf-st: Video captioning based on key semantic frames extraction and spatio-temporal attention mechanism.2020 International Wireless Communications and Mobile Computing (IWCMC), pp.

Augmented Vision-Language Models: A Systematic Review Ksf-st: Video captioning based on key semantic frames extraction and spatio-temporal attention mechanism.2020 International Wireless Communications and Mobile Computing (IWCMC), pp

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:39.637323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:39.637323Z digest=sha256:6ac76406ac98678818f26561eb523cbb5d849d08d807c5e9d35e34b96454f391

Observation 10b1a622-8177-4cc3-9efa-1a773ebc9c29 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Augmented Vision-Language Models: A Systematic Review Learning Transferable Visual Models From Natural Language Supervision

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:39.833296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:39.833296Z digest=sha256:0de6b8156b0a51c9d033d2d78630f949c51db0d5f3ac09aeecc4e31b1b29927f

Observation cb293b46-f70b-41e0-b0d5-e7d651d513d1 · outbound

This paper cites VLC-BERT: Visual Question Answering with Contextualized Commonsense Knowledge.

Augmented Vision-Language Models: A Systematic Review VLC-BERT: Visual Question Answering with Contextualized Commonsense Knowledge

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:48.677769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:39.912968Z digest=sha256:cdae2f9f22311d86ab6c54ae46a00ac85c58300d1e456fcd466211b4762f158c

Observation 74d88e8a-e7c2-466c-994b-71125f6ddc2d · outbound

This paper cites Outside knowledge visual question answering version 2.0.ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.

Augmented Vision-Language Models: A Systematic Review Outside knowledge visual question answering version 2.0.ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:39.993751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:39.993751Z digest=sha256:6fc2a9013e8a9d77c64d696efab54fc02e51ba31942f72c12e8eeeb9a6dbf706

Observation 5f5c5dad-f8e2-4d82-9629-d6530af0c350 · outbound

This paper cites "Why Should I Trust You?": Explaining the Predictions of Any Classifier.

Augmented Vision-Language Models: A Systematic Review "Why Should I Trust You?": Explaining the Predictions of Any Classifier

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:40.096217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:40.096217Z digest=sha256:15a670fd4736d8d69e8c820928b97aa02d38c8ba0e330bcc9b1e290222d296e9

Observation 833c3f63-3e79-41f5-91e8-d66e8a42f87a · outbound

This paper cites Interpretable Machine Learning: Fundamental Principles and 10 Grand Challenges.

Augmented Vision-Language Models: A Systematic Review Interpretable Machine Learning: Fundamental Principles and 10 Grand Challenges

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:40.159564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:40.159564Z digest=sha256:46e5d1710df8eed91142aeb5e8a9dcb7832d6dc2ee8a62e824eea9e63d991a28

Observation e7a7a7a0-e4e4-4fe5-a56e-7885a52dc24d · outbound

This paper cites Alireza Salemi, Juan Altmayer Pizzorno, and Hamed Zamani.

Augmented Vision-Language Models: A Systematic Review Alireza Salemi, Juan Altmayer Pizzorno, and Hamed Zamani

Reference 95

Resolution
verified exact
doi, observed 2026-08-06T14:33:45.513542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:40.311849Z digest=sha256:f8ee77f2c7fe5979bd12529786e665315b70313f3723c7c6a3a0698e89bbaaf9

Observation e3b07021-ac50-4269-a91b-1c9318ee16bd · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

Augmented Vision-Language Models: A Systematic Review Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:40.468859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:40.468859Z digest=sha256:249304aa1846e1882cab6fcf7f8971bb3905376c0f1ed4d0c291c11dbe735c3c

Observation 955cbaf7-2781-4c65-98bc-e04ebfd96c63 · outbound

This paper cites Graph Neural Networks in Vision-Language Image Understanding: A Survey.

Augmented Vision-Language Models: A Systematic Review Graph Neural Networks in Vision-Language Image Understanding: A Survey

Reference 97

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:33:48.511511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:40.598922Z digest=sha256:4d70e5b7f96eecc3a42765948c6ff706695dce5ca0b562a74be9333a0c022d35

Observation bcb998fe-bf02-441c-81f9-7e249386b5e4 · outbound

This paper cites Logic Tensor Networks: Deep Learning and Logical Reasoning from Data and Knowledge.

Augmented Vision-Language Models: A Systematic Review Logic Tensor Networks: Deep Learning and Logical Reasoning from Data and Knowledge

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:40.743779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:40.743779Z digest=sha256:b34f6f19d932ad7c4c8ef01ba9ce45981f280f662725d6a24bc9bc17964c652e

Observation caa90521-d631-4450-abe6-cb924d4d8a5f · outbound

This paper cites UniRAG: Universal Retrieval Augmentation for Large Vision Language Models.

Augmented Vision-Language Models: A Systematic Review UniRAG: Universal Retrieval Augmentation for Large Vision Language Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:40.930906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:40.930906Z digest=sha256:b1afb886c307d34cc00b9192474d737302c655364af8c930aa7e20546693023a

Observation a638ccbd-41f1-43ea-bb56-70ca68babd8e · outbound

This paper cites VCD: A Dataset for Visual Commonsense Discovery in Images.

Augmented Vision-Language Models: A Systematic Review VCD: A Dataset for Visual Commonsense Discovery in Images

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:48.373864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:41.024746Z digest=sha256:3911cfb68cf0cd1dacc392556f25129c1cd80f6ffd0516e960fb7f579599c584

Observation 2ce01051-3615-4af2-a21b-f37b9eceee81 · outbound

This paper cites Reasoning over Vision and Language: Exploring the Benefits of Supplemental Knowledge.

Augmented Vision-Language Models: A Systematic Review Reasoning over Vision and Language: Exploring the Benefits of Supplemental Knowledge

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:48.201075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:41.174422Z digest=sha256:e0f2005a13fb7e2429da61425bb3dc25926c362ffd29f06967b6632f68af1c8a

Observation 5b38d500-9f46-473f-84cb-573ed274abcc · outbound

This paper cites an unresolved cited work.

Augmented Vision-Language Models: A Systematic Review Unresolved cited work

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:41.264804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:41.264804Z digest=sha256:875084bc1ed94480f12b60e39739ea144a0a11c6d9f24ebfca12359e90879fbb

Observation 1003a2c8-0f9a-4070-b715-a8eb6d8805a3 · outbound

This paper cites Singh, Anand Mishra, Shashank Shekhar, and Anirban Chakraborty.

Augmented Vision-Language Models: A Systematic Review Singh, Anand Mishra, Shashank Shekhar, and Anirban Chakraborty

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:41.378191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:41.378191Z digest=sha256:a0efd85978c2c82d05336155b0f5fef46b7d0c78ca6553a1c24f09c49de30b35

Observation ba1a456c-e53a-474f-b604-ce30a00a0bc0 · outbound

This paper cites Liu, Yang Yang, Xuequn Shang, and Mingxuan Sun.

Augmented Vision-Language Models: A Systematic Review Liu, Yang Yang, Xuequn Shang, and Mingxuan Sun

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:41.423630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:41.423630Z digest=sha256:936ed1631912412b67ad298f15ba4401bc974fad5a2b1ed6470e5dc7ebcd9b47

Observation 5e4337c5-85f7-4126-b9a1-43b3cf67efda · outbound

This paper cites ConceptNet 5.5: An Open Multilingual Graph of General Knowledge.

Augmented Vision-Language Models: A Systematic Review ConceptNet 5.5: An Open Multilingual Graph of General Knowledge

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:41.533386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:41.533386Z digest=sha256:6b0ff13f0b4f478f8cd3caae26c1f1b55e14ee6f51dfda5051b87aa5585edbc0

Observation 170791cf-d1ed-443a-ad18-a75679311381 · outbound

This paper cites SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs.

Augmented Vision-Language Models: A Systematic Review SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs

Reference 106

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:47.991462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:41.656102Z digest=sha256:f003aeac0b4a28a37e86c340bba151a415d67c30cc3ad0b6c6bf947f4de23254

Observation 35454628-c8b7-44fb-9538-688915989f08 · outbound

This paper cites Learning Visual Knowledge Memory Networks for Visual Question Answering.

Augmented Vision-Language Models: A Systematic Review Learning Visual Knowledge Memory Networks for Visual Question Answering

Reference 107

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:47.868348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T14:33:41.772208Z digest=sha256:bd4a38ca62e9be4cf3a94aeb51e90530f9dc6203452bcb7d5f6106eddadc20f1

Observation 43fc39b9-9292-422a-b382-b8457345f13b · outbound

This paper cites Modular Visual Question Answering via Code Generation.

Augmented Vision-Language Models: A Systematic Review Modular Visual Question Answering via Code Generation

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:41.947610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:41.947610Z digest=sha256:10605a49599df43fe1facf59b0af8ea73e8e82eaaaa5dc68340e311d4b6198d7

Observation baafa8ad-acf8-48ae-84ba-c116dbbfbdb9 · outbound

This paper cites ViperGPT: Visual Inference via Python Execution for Reasoning.

Augmented Vision-Language Models: A Systematic Review ViperGPT: Visual Inference via Python Execution for Reasoning

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:42.097727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:42.097727Z digest=sha256:156b1e720863bc7471e63b6115cede7568d40e0a007f5ad9308de6cfc322fcb8

Pith citing papers

No inbound Pith citation observations are available.