Pith. sign in

Paper Citation Record · LEDGER

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

As of 13 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 16 inbound Pith citation observations for arXiv:2411.15041.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15041 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:38:17.374945Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:15:14.692955Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T12:46:14.749364Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a3c20d6e-0cef-44fb-9997-70b5f8f5731a · outbound

This paper cites GPT-4 Technical Report.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.146461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.146461Z digest=sha256:1d8d58a0338e15b36f001f1615ae012952a8f13197bce84ecadae1b2a27d3ceb

Observation f3e98cd8-a5fa-430a-a368-41bbf5c112d8 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.151337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.151337Z digest=sha256:1732906514211db97e340fd8b916c1ca7039c78cd35efcef16db9a060edc8d6b

Observation ce40568f-fb40-46f3-bd33-f3628fb91f7b · outbound

This paper cites How (not) to ensemble lvlms for vqa.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA How (not) to ensemble lvlms for vqa

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:18.085712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.155855Z digest=sha256:5908f0cadab1a1b41bf50768565453ec901cdfb06a2949be4b7e822d980dfe1d

Observation a9cd680d-0338-445a-b3c8-20ec4459d95c · outbound

This paper cites Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.160452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.160452Z digest=sha256:9b7d6bb7fe01c8e5e4bba193a3f580cfbc68dee0aa2125dbab52c006fb482366

Observation 55e2dce2-3ece-4590-a49a-b6e76ad78783 · outbound

This paper cites Language models are few-shot learners.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Language models are few-shot learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.165435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.165435Z digest=sha256:ceb5939da5892aadf8c3f654dd9dedd271c1e238a6007555459eeb0e83c4f9b8

Observation ab987f6a-2223-47f6-87a1-8aa7816ad245 · outbound

This paper cites Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.169982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.169982Z digest=sha256:1b0fb77f3841452e93cacefe3031fadd618d19a60ffce1b8f80db8841a4f6af8

Observation f8bd7257-79aa-49d1-8937-302667a75f10 · outbound

This paper cites Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.174603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.174603Z digest=sha256:965c52f2ac549dd72956671f9c0cfa6e346ff781810a29d14181350c42d457d8

Observation 05816e27-ebaf-40ab-bc50-8651b3e077f3 · outbound

This paper cites HAMMR: HierArchical MultiModal React agents for generic VQA.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA HAMMR: HierArchical MultiModal React agents for generic VQA

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.179802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.179802Z digest=sha256:8388b75bd2cb41be420c8cdec71c469b5c3ba39a0d47ee9b20033a6e497d5d4d

Observation 4456f969-36bb-42d6-bff1-e8eb665d2c03 · outbound

This paper cites an unresolved cited work.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:18.062083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.184256Z digest=sha256:83abaf5d664c8a319d510f3e4e1c8b6ba4e086c7870759ce6e32789e156c1a3d

Observation f5b49162-afcd-44d6-a54e-7528494e9e16 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.188318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.188318Z digest=sha256:2c6668e20850cc1f0f28db21987a40d5920ab5c66e8b930ee5df1e5e5f66b3a6

Observation 8f572509-f4cc-4bf5-b363-6fa4b24d8456 · outbound

This paper cites Palm: Scaling language modeling with pathways.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Palm: Scaling language modeling with pathways

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:18.038162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.192424Z digest=sha256:b51d7d39df2a816847c02015dacc16e6344114b9f858f9e6b085d1cfbd04bbde

Observation 30a3ebf1-d3f3-45af-a0d6-408d6535ba4e · outbound

This paper cites Scaling instruction- finetuned language models.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Scaling instruction- finetuned language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.196574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.196574Z digest=sha256:13007f78530c420f1aeb751722befc2097e3bb1fce74c6db8c76b6e9d053b89f

Observation 673e03eb-afc0-4d32-8506-58860c4bf8e4 · outbound

This paper cites Mme: A comprehen- sive evaluation benchmark for multimodal large language models, 2024.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Mme: A comprehen- sive evaluation benchmark for multimodal large language models, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.200643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.200643Z digest=sha256:297c7d9883c5f9276eef2686ee45c355946bfcbdeba8388aadd4eb49eae1494b

Observation aba424c8-e18c-499d-918c-359ad92bb8e5 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.204767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.204767Z digest=sha256:8ab768887bc90a67929a8d1ebb0f3d5e7a7dabf9e4ee5e830fccee63b4c02ca8

Observation 95a2d3b3-ba18-42eb-828b-2ba12a776ae3 · outbound

This paper cites Unsupervised dense information retrieval with contrastive learning.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Unsupervised dense information retrieval with contrastive learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:18.005987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.209217Z digest=sha256:a29b8a61faf2069a2ccc5397034578ab8bf52587484c46eb222e97c36544b0dd

Observation 1c128526-bf95-489a-b750-8f142da0b5de · outbound

This paper cites Google Lens.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Google Lens

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.992389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.213191Z digest=sha256:0b3a3ee49c435301bef6e8cd721a66dc00620151441fb4eba3bba767c79cc38c

Observation adf40151-1bb2-483e-bd5e-e5fda5a8960f · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.979117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.217344Z digest=sha256:409b1dcc936cf09358549c504aa9ec004fcfbf729f931b0fa3ca5f959e9aa5d7

Observation 09d5a0c2-3288-431e-9729-1e09661e371a · outbound

This paper cites Avis: Autonomous visual information seeking with large language model agent.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Avis: Autonomous visual information seeking with large language model agent

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.964047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.221519Z digest=sha256:d3b5bedb9123774a405d41a9fcc1e57edf3fdf77e720d679d78c6669a6efe7a8

Observation 2a4e55f7-27bf-4cec-88b8-353650d42865 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.949606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.226196Z digest=sha256:6246e118ebc6113193668ab8e6b02c80b3938913cc041fc4f36fe6f056304abf

Observation f1fe4afb-78e8-477a-b052-f863e3ac0d39 · outbound

This paper cites Large language models know what is key visual entity: An llm-assisted multi- modal retrieval for vqa.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Large language models know what is key visual entity: An llm-assisted multi- modal retrieval for vqa

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.934152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.230295Z digest=sha256:42c2a5137d8ef44350b3a5c004511a0a0349d2e2ddeb966da162fe40d77721f5

Observation 730f4c11-163c-4f86-b3fe-2ccf551349d7 · outbound

This paper cites Large language models struggle to learn long-tail knowledge.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Large language models struggle to learn long-tail knowledge

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.920445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.234502Z digest=sha256:358fb700c6d2b2870b009cad299cc802e19f9efaa6561575a9371af55521f0fc

Observation 3e1c9ab7-9076-457c-9abd-6045024c5b39 · outbound

This paper cites Natu- ral questions: a benchmark for question answering research.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Natu- ral questions: a benchmark for question answering research

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.906411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.238466Z digest=sha256:edf3f24e50a54a24310850edbb147c68396c88e2aad0fdd00369ff5d557e16e8

Observation fab9994a-ec10-4510-82b2-1b3bde0813e1 · outbound

This paper cites Viquae, a dataset for knowledge-based vi- sual question answering about named entities.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Viquae, a dataset for knowledge-based vi- sual question answering about named entities

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.892269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.242588Z digest=sha256:17421078f2b613169d84f92003d8e25c3566ba6c15dd5a1c1c0ea853cc3af228

Observation eb30ec1a-a9ad-4225-98b7-bea6d7090a45 · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Retrieval- augmented generation for knowledge-intensive nlp tasks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.878452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.246759Z digest=sha256:d3c657b62d40c372a4ded338b59a7726b7f0ea458cea1b0c1a104f0649eafd7a

Observation 33fa80b8-2a9c-44a4-abf5-b9e03dba5dfe · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.250707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.250707Z digest=sha256:a5de3bacdb73f3e216cec67bbe08fd6332b9585ef49203e29dc94cfda7a8c36a

Observation 75c1788c-a275-45d5-93fd-5f795fc7ac9c · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.254872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.254872Z digest=sha256:253c4199c8d44654a5709cee62070167b119ea38e0b40d70c4a3925583105682

Observation 82d6670f-8a86-46e4-ac20-801d53503970 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Evaluating Object Hallucination in Large Vision-Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.259644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.259644Z digest=sha256:a475ba95949586a748ff4cff760cd0ac262272a714fbf6be15db46827577ddc9

Observation d61a41e1-dd37-47fd-9fc7-d72a0ec50854 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.264576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.264576Z digest=sha256:2254d89331cdb44a9ec26cd86c0a217315117d2a798b89d98f4b4d29a391b721

Observation 94f3a57e-df3f-4218-bff7-1ec5b6f6ec63 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Improved Baselines with Visual Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.269363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.269363Z digest=sha256:1e1de8fa04ab6cbb568cfd6125077a635d5febf4aa8ae273e6859e62f44af603

Observation 07b85807-d788-4967-bce3-aa91d8aeec63 · outbound

This paper cites Visual instruction tuning.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Visual instruction tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.855345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.274343Z digest=sha256:da00345ae7281504dab95e4f0fb3c48c138eb200b39b7042daccc1898b397c39

Observation a67bb6ea-8908-489e-8bb8-49f90053ca85 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?, 2024.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Mmbench: Is your multi-modal model an all-around player?, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.278346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.278346Z digest=sha256:7b07a9b882606da0cdb40f8c4166e2f42f52c31f7875d1a267639ba4f9151a53

Observation 64c369a7-86bc-4731-8268-e636b720a62c · outbound

This paper cites Generation-Augmented Retrieval for Open-domain Question Answering.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Generation-Augmented Retrieval for Open-domain Question Answering

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.282351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.282351Z digest=sha256:097154d7377174fbb4503c0294600dcdedec90bf76449c4c35a7e090bf34f7c7

Observation 4ec920ac-0f10-4557-a05b-cc7baf6d880d · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.286631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.286631Z digest=sha256:46365416a9f69b28c8c0b8ac7d5a627a21dc7dd1d2e9f8630db52b4a070fe98b

Observation d6b2b1a5-9b93-4f81-828a-6ba4f13f0e2d · outbound

This paper cites Sfrembedding-mistral: en- hance text retrieval with transfer learning.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Sfrembedding-mistral: en- hance text retrieval with transfer learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.823573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.290736Z digest=sha256:edd87f02fcc4fba99dd83c7831a18dbeb777f1c247208710ec2a78875679a85f

Observation 173500e1-cc88-4fdf-8289-2ebb48f75c30 · outbound

This paper cites Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.810059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.295096Z digest=sha256:deac7358976d0f82d08a069ab0e224d9e12fdbcf7d3904078ef24648f2d079a6

Observation dc030e7e-2e92-4eb0-b906-5a824726877c · outbound

This paper cites Plotqa: Reasoning over scientific plots.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Plotqa: Reasoning over scientific plots

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.796082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.299590Z digest=sha256:e323c391bcdce432bb8a43d1a4fdbef2f360a2d9c8463bdd9ae7f951c4d1b015

Observation 737fee84-f23b-4adc-9cc4-9e3f14310b4c · outbound

This paper cites Introducing gpt-4o: Openai’s new flagship multi- modal model now in preview on azure.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Introducing gpt-4o: Openai’s new flagship multi- modal model now in preview on azure

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.782797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.303681Z digest=sha256:42cc230569bb6245f68445f06ca40fdd6aa5fb30b7575455bc6ec89da5d5ae00

Observation c966edb7-4ef3-4e87-823e-60358f81d130 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744,.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.312457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.312457Z digest=sha256:b503f791f031b9ba9f226a1ee636fb57802e14b8765d16a0fd8468691a6f4b3f

Observation e46fa1f1-e8bf-43cd-a5d3-efb44235add8 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Learning transferable visual models from natural language supervi- sion

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.316597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.316597Z digest=sha256:97f3819e2c8e766017f35c3f22761995102fdffd314599bfbea3f52435d2dffe

Observation 2b6fbaac-fb38-44fc-b880-d4029cd38b65 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.320542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.320542Z digest=sha256:96c2cc3f6c69fdda8e4dc074bddef2a2f888cf6a280f00a4498031b71d7f976a

Observation 42695280-89aa-46a7-8183-ddad4e68803b · outbound

This paper cites A-okvqa: A bench- mark for visual question answering using world knowledge.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA A-okvqa: A bench- mark for visual question answering using world knowledge

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.737761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.324856Z digest=sha256:eecb3e79eb61cc5b8e2c110e4f696a0f78d2e1e0ed7d8f5716431c8c06a04cb8

Observation 4e0e30aa-0bd6-48ae-a0fa-86515c51ca48 · outbound

This paper cites 10 Towards vqa models that can read.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA 10 Towards vqa models that can read

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.724361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.329098Z digest=sha256:3d284310d36d74bd78de046f3e3c52dee4d9132da4685d06f9e330b38cacc05f

Observation cc427235-2777-4bb9-92cf-7607e25faa02 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.333113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.333113Z digest=sha256:6fc6749a616a2a3cf843832d9aacb135187264781d87001366c371841c2a947e

Observation af0c4a4e-6778-470a-98fe-4985de76c81f · outbound

This paper cites Chain-of- thought prompting elicits reasoning in large language models.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Chain-of- thought prompting elicits reasoning in large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.337437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.337437Z digest=sha256:c5c2f59f57f2586416d568cfa960d3b8e4dfe47b7a71dccee0ee24509d481161

Observation 146dde20-b8bf-41c3-89fc-f4edabb449a5 · outbound

This paper cites EchoSight: Advancing Visual-Language Models with Wiki Knowledge.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.341496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.341496Z digest=sha256:78c6c78abd2a8d9bcef0ea2255e9fb8d673a94b412b2f8801d6277c81f856104

Observation b120ab8b-6f06-4302-8517-6a6b428ff9fd · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.345942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.345942Z digest=sha256:2c0119ac6f33d0c2baeb31bd13dad1b4aeac9846327314c578f4f38f06a63c10

Observation a10c10ec-bcec-43e6-86e9-a4a05af3e670 · outbound

This paper cites Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:17.350239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:17.350239Z digest=sha256:bcbcfb968d862816e1a73b6498981f07269820d9e49aa51f582b4324b12dca58

Observation 33ba0bbd-e8b4-4299-ae6a-7c116f5814f7 · outbound

This paper cites Mipha: A comprehensive overhaul of multimodal assistant with small language models.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Mipha: A comprehensive overhaul of multimodal assistant with small language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.701236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.354552Z digest=sha256:fdc93b31169435ee2cdb1da2cbfd76e81937ad4d9b3bd0485ba28f43788ebddf

Observation 513aa7f6-e4a0-4a5b-8802-95233b90f341 · outbound

This paper cites an unresolved cited work.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:17.686628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.358544Z digest=sha256:b6256b8bff3c208ecddd8c98c38a0bf6e1b8f72b6889d25ecf4b6309b184787c

Observation 15f19c85-2bb9-40f4-8df8-6d3cc4ba2661 · outbound

This paper cites an unresolved cited work.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:17.672710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.362733Z digest=sha256:53d626cff7b75b16ed8746e281042eefb649a28a8f202ae0e21c0c4c8a192a8a

Observation 8c34f509-06b4-4e36-b36e-383fb428a735 · outbound

This paper cites Answer: {answer}.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Answer: {answer}

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.658988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.366776Z digest=sha256:107fa6022301e3fe579730acff94e62cd069a00bdc6a037ea15796d09ba81303

Observation 4301decd-d900-47f6-b12e-3248c3fe3575 · outbound

This paper cites In the without external knowledge setting, the model relies solely on the knowledge encoded in its parameters to answer ques- tions.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA In the without external knowledge setting, the model relies solely on the knowledge encoded in its parameters to answer ques- tions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:17.644317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.370966Z digest=sha256:ebb85c8dbfff8ff74d23151db83f7c0b640cdc31450e76c9869d648f87218f10

Observation 363e23a3-a50b-4f1b-bb68-8e4e2b6cc397 · outbound

This paper cites an unresolved cited work.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Unresolved cited work

Reference 54

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T14:38:17.629388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.374945Z digest=sha256:7f37d76cfe4ab56afdc57a2647a72fa482ee856698a4e83ba79de36885c62a9c

Observation e7b48f32-9350-49a3-915b-251fc9205009 · outbound

This paper cites an unresolved cited work.

mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:17.769315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:17.307785Z digest=sha256:1bfe0b6df9d778de5da058b5f59f533489d6b77620d05f03ef4f3201a8c40b28

Pith citing papers

Observation 571d8895-cb81-4a6f-9f88-fd0577c18585 · inbound

Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation cites this paper.

Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:52:19.993606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-19T13:50:30.090068Z digest=sha256:a279dc4e403a9d9db5dfd4a7c9ec878325312096b4b592f03dec168167b29b58

Observation 875aafb3-c194-4219-98e1-50d24b64f030 · inbound

CoRe-MMRAG: Cross-Source Knowledge Reconciliation for Multimodal RAG cites this paper.

CoRe-MMRAG: Cross-Source Knowledge Reconciliation for Multimodal RAG mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:42.371621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:27:42.371621Z digest=sha256:ac31bcc0128d8531af2dd9fa09448e82172718c0f7c824f82f2b7e6e02536f0b

Observation 0028bbce-5de7-4a43-97ad-17a3629fb18e · inbound

mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQA cites this paper.

mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQA mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:06:55.031490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-19T00:05:08.866244Z digest=sha256:0fd8f9ed3e09f83207091da98d70a02e62a4a8601c1c3dc950b777cf1d1a373e

Observation ba192c7a-0448-4bcb-b339-ddc3dd8396d8 · inbound

Towards integrated sensors for optimized OCT with undetected photons cites this paper.

Towards integrated sensors for optimized OCT with undetected photons mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T23:26:28.640658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:26:28.640658Z digest=sha256:c7bded32d4e56314f70e0024128d11d87e5119900bb904f161601def767a3eb5

Observation b2bdfe73-d43b-41af-85ab-ae2ccc4bbc60 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 111

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:53.107445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:53.107445Z digest=sha256:35713783606167f6f1e5667ef07d93f9cf4f9b7d1691a2c24a8d6ed2d98a0f1b

Observation 1cef374f-1c2e-4445-9c40-1aae5369ac16 · inbound

Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering cites this paper.

Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:06:49.834243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-18T20:04:52.852253Z digest=sha256:83095857d981692d41adbeb15fc7405f153dc5bb65a40430967712ef3da978de

Observation c3f50af2-666a-4c96-9b4f-03dc62a0dc58 · inbound

Recurrence Meets Transformers for Universal Multimodal Retrieval cites this paper.

Recurrence Meets Transformers for Universal Multimodal Retrieval mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T20:06:09.182830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:06:09.182830Z digest=sha256:b5097df99d7bdcdc182656f081b4e9dccfafee3817a36113401e29105b32e535

Observation 75576e0d-ca0b-4215-afda-05b752adcfb2 · inbound

QKVQA: Question-Focused Filtering for Knowledge-based VQA cites this paper.

QKVQA: Question-Focused Filtering for Knowledge-based VQA mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:54.028617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T12:53:59.668663Z digest=sha256:5369afb7c148786dc95a3b4e48388b8f1db869648a1799b54626c39932baef0f

Observation cb79a4e0-3510-432c-8642-b6680e020dfe · inbound

WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering cites this paper.

WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:48.065154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T19:59:10.657346Z digest=sha256:d47eb7647a37b3aa2e85d2f2df516bc68736d818e0c886a83410407f7f9ed9c8

Observation 8cb9be3d-a963-4268-b521-4e1250a7b597 · inbound

CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning cites this paper.

CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.730102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T08:14:20.558527Z digest=sha256:b3864b36644d3940508bb153989236350d4ef27e87346339b5ab9e8329b6c6de

Observation c2c2eb77-cb22-404b-8831-8d3126503c4c · inbound

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification cites this paper.

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:59:46.839046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-26T08:12:14.829556Z digest=sha256:4396e915c4a9c92af142cff1dab13bc1f6a2e54a75019547be6939da1bb4f758

Observation aa6dbd07-ebad-41b7-b8e4-c04f6228adb9 · inbound

MMAgent-R$^2$: Learning to Rerank and Reject for Agentic mRAG cites this paper.

MMAgent-R$^2$: Learning to Rerank and Reject for Agentic mRAG mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T12:46:14.750625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-09T12:37:07.906342Z digest=sha256:f2f4bb2c268d753412c6220d5d5a742798c8e28dedf84341350bb0ccc53c670f

Observation 11576698-4461-4e44-88f8-c8736543b8bb · inbound

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG cites this paper.

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T10:20:50.701345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:20:50.701345Z digest=sha256:48a1fa20121d105587690689d88561266ddd1c80f8cec70161f93e875f527445

Observation b8c9e70d-aa80-4e22-a27a-d2696ec8f8d1 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:17.751661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:17.751661Z digest=sha256:92b11102ed69eb2613367718b5ef9b2e1f5fef4c7e73bb32f73e3e4bd3109c4d

Observation 90a75562-a2e8-494d-8a15-9392fa7bb123 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:53.771485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:53.771485Z digest=sha256:b665ea7cb667c854106022e4247647737729074263d84f572fe467a9a8d0ca8c

Observation 5069e662-c10a-44cf-aef2-4d67d23e3bcf · inbound

M$^3$Prune: Hierarchical Collaborative Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation cites this paper.

M$^3$Prune: Hierarchical Collaborative Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T20:15:14.692955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:15:14.692955Z digest=sha256:85da4475684fa8759bed5d3848b3ea99396466eae205ed81106ae2737b06e4dd