Pith. sign in

Paper Citation Record · LEDGER

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems

As of 21 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 3 inbound Pith citation observations for arXiv:2605.19307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19307 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T07:08:27.096029Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:34:36.466548Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T22:17:26.215530Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact13
  • verified fuzzy31
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 40a77528-40e2-4d85-a656-377706929da9 · outbound

This paper cites StoryLLaV A: Enhancing visual storytelling with multi-modal large language models.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems StoryLLaV A: Enhancing visual storytelling with multi-modal large language models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.630099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:c0056d678ed2c0527703bb8e0587bc52a60dbfd42c38400193a43899522b4e59

Observation a32790f2-14ba-4cbe-a9c0-7e9a045cc975 · outbound

This paper cites Refined semantic enhancement towards frequency diffusion for video captioning.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Refined semantic enhancement towards frequency diffusion for video captioning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.634289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:d1c5822161f67eadd72a7e1e9bca3ea9e7041a2032f8ea4d204462a66af7103f

Observation e0a285ca-b2af-408a-9789-2ffe8a79dd31 · outbound

This paper cites Action-aware linguistic skeleton optimization network for non-autoregressive video captioning.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Action-aware linguistic skeleton optimization network for non-autoregressive video captioning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.615134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:53f3081bef18dec2663f6813089dd2d3076686476f0411c82f860f115f269452

Observation 2cb57102-b457-4f1f-afab-e295cbfff7bb · outbound

This paper cites VQA: Visual question answering.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems VQA: Visual question answering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.637677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:0d83f3db2684b518818c1e039ccc05fdde0e61faeb74b045963b9e15d68a690c

Observation 366adef9-3c3d-4ed0-9bff-125b5d6fcf40 · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in visual question answering.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Making the V in VQA matter: Elevating the role of image understanding in visual question answering

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.659541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:9a6f5d608837b9a02032382e6bbc634ab98de5605366b44df4a08fdfa5999319

Observation b8403d03-2906-4045-9077-0d8bc97bb874 · outbound

This paper cites Robust visual question answering: Datasets, methods, and future challenges.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Robust visual question answering: Datasets, methods, and future challenges

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.648452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:1966963f190b3db3dbac48161556cc6ce9bdf865fe61f25df49a6eee0617376c

Observation 9b409ebb-1613-4226-8bee-c783736fb4ae · outbound

This paper cites Metamorphic Testing: A New Approach for Generating Next Test Cases.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Metamorphic Testing: A New Approach for Generating Next Test Cases

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:13:06.636646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:68f99a60dd79367275468c34bcd0771a4971c085320c89a464f59b2275379e8f

Observation a3cc4d09-97ab-47bc-bca5-3d667ce72d19 · outbound

This paper cites KVQA: Knowledge- aware visual question answering.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems KVQA: Knowledge- aware visual question answering

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.619540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:043cd794424a49ad58f907b56772be6105ee82c70f3566cb8bcf35ef5c0d73fc

Observation 7b58c1b2-cab4-4b3d-8450-7f94e5fc5db4 · outbound

This paper cites OCR-VQA: Visual question answering by reading text in images.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems OCR-VQA: Visual question answering by reading text in images

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.613066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:91b6be16c0467f94aa7ff21d47123fd4cb60f14e9e635b190df87c7bdc98986a

Observation 5a800ab6-245b-4278-a216-2196e09496fe · outbound

This paper cites Metamorphic testing: A review of challenges and opportunities.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Metamorphic testing: A review of challenges and opportunities

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.635651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:9cd47b50b3c19214e9d594d01024082eb9ed802851b26a85bc5210db6a241c13

Observation eba38e50-7dcd-4ad7-a71b-1bc60426b74b · outbound

This paper cites Perception matters: Detecting perception failures of VQA models using metamorphic testing.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Perception matters: Detecting perception failures of VQA models using metamorphic testing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.655114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:8a47613288c91d1abd2c5d8f92b4cd82d91867de4dc43b0b7885e852cb01e493

Observation b1ea812e-da81-412c-8d65-ec0f9b9bbedd · outbound

This paper cites Metamorphic testing of image captioning systems via image-level reduction.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Metamorphic testing of image captioning systems via image-level reduction

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.617226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:613c7e90c98c87146415af90294a6cb19d1fb93bc33cc4d5b0267398ca8fae07

Observation defa4c4a-2175-4611-ba56-e3c7800fc10c · outbound

This paper cites How Multi-Modal LLMs Reshape Visual Deep Learning Testing? A Comprehensive Study Through the Lens of Image Mutation.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems How Multi-Modal LLMs Reshape Visual Deep Learning Testing? A Comprehensive Study Through the Lens of Image Mutation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:13:06.627458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:0a0ced870d5683e4479441af05bbe8a3ef39adb02d47a418545594d3a5a6fc28

Observation b2b34d11-a9fa-41aa-b0ab-9409456710bf · outbound

This paper cites CLIP in mirror: Disentangling text from visual images through reflection.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems CLIP in mirror: Disentangling text from visual images through reflection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.638490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:65a03ba689c5c2682e32f880a40ad95b984e4e3a2df367099d2ca21d1f20fc4c

Observation dc90bcf0-16f3-4835-91ad-7c23000b50c4 · outbound

This paper cites Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:13:06.643279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:de452d51969e506e13c4a9e30eb11c4a29ccf387a529507ede7f8dfebc4fca08

Observation 9cad04b3-a962-45c0-b40b-1064c7646f41 · outbound

This paper cites Improved baselines with visual instruction tuning.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Improved baselines with visual instruction tuning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.608107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:d6940e3031322a2750d4984c6c5c2cad43aa09ac42943dab1d856db76364eb65

Observation d17fc3ed-8c8e-4869-ab75-3499af4d8a91 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:13:06.658617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:e9c0451bc540dc465286f3cce575aa53afa09855973fe2ac2d2c7f358caecd66

Observation 7d250587-ca1f-4825-ba6a-0dc71fcb901f · outbound

This paper cites Qwen2.5-VL Technical Report.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Qwen2.5-VL Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:13:06.633393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:775e3d928a40d7162c9001cbed57086c1276156c5833c12a92c99eeca4a43975

Observation f4dfff8c-3e48-48d0-9efa-f8f3188e200c · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:13:06.624537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:c8fac9c883a31811e885cbcc84b03017bdbe635400d3d0c47540407061815f57

Observation 3f4fcffa-d64c-4203-9598-e9edc0a06f77 · outbound

This paper cites Cross-modal retrieval for knowledge-based visual question answering.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Cross-modal retrieval for knowledge-based visual question answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.651294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:dc4151e0554fd4f7a65259e03c1c132da4a8d341a0130cc7b9ac0350456df6fa

Observation 7299ff26-4b25-4e92-934d-911836cd0826 · outbound

This paper cites RoRA-VLM: Robust Retrieval-Augmented Vision Language Models.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems RoRA-VLM: Robust Retrieval-Augmented Vision Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:13:06.630472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:7fdf639e6a1184a4699907cafeddd9d0a6fb20f888c6fba4bd796d6851a484f5

Observation e8c68954-cbd3-404f-b165-107084698c7f · outbound

This paper cites EchoSight: Advancing visual-language models with wiki knowledge.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems EchoSight: Advancing visual-language models with wiki knowledge

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.601543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:ac6003edbabbb493fe6198afdadaf8722f9658e90d1426e1a3e00d391dde7ecf

Observation 9863a52e-f48c-4a4f-be88-930f525f8f76 · outbound

This paper cites Wiki-LLaV A: Hierarchical retrieval-augmented generation for multimodal LLMs.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Wiki-LLaV A: Hierarchical retrieval-augmented generation for multimodal LLMs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.628216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:043c0c9184c03d171b8939357d4a4f20bde69b402588b9514d9db9ffcbc05b98

Observation 9d401e5e-077c-4c7f-b864-2c7217f1b398 · outbound

This paper cites Augmenting multimodal LLMs with self-reflective tokens for knowledge- based visual question answering.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Augmenting multimodal LLMs with self-reflective tokens for knowledge- based visual question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.608935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:d8c5ddf04de8bc6698d0878b32986d3af6d1474391761459d513ecd022a05bf6

Observation 5259ee6c-ea88-4d77-949b-dbb3da6fd157 · outbound

This paper cites Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:13:06.621722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:afbdb3efcc020052f945137b521ea40636fcee4fdf38a5c5ae5670aa4e25655b

Observation 9537564a-1628-4075-be86-33cee8c52a11 · outbound

This paper cites MMKB-RAG: A Multi-Modal Knowledge-Based Retrieval-Augmented Generation Framework.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems MMKB-RAG: A Multi-Modal Knowledge-Based Retrieval-Augmented Generation Framework

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:13:06.649910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:1eef40138f9424eeac57fc3d24072f09d1212b1985aa96020b8a049a82225cfc

Observation 13892cdb-6abd-4c60-a392-2b935ece5dc9 · outbound

This paper cites Knowledge-based visual question answering with multimodal processing, retrieval, and filtering.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Knowledge-based visual question answering with multimodal processing, retrieval, and filtering

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:13:06.639865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:454554b515f3cc7afb1a4e2df7b36991e68992842bd6956ef0d2cb3b925652f4

Observation b93bf570-9e8d-4a8e-a89b-100242817fee · outbound

This paper cites Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.636496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:cd9cd24fd8487694aaf126adf9dbb4cc90a53a93434f3beb672ac6cc7a3c3d6b

Observation 23acc7a2-4cb9-4009-8d5d-e0f358c54932 · outbound

This paper cites Can pre-trained vision and language models answer visual information- seeking questions?.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Can pre-trained vision and language models answer visual information- seeking questions?

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.653221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:c87cb70d7a358f0a5d3bfd1efe7623492743af645468bffcc5ac065d7c15e010

Observation a9766a17-b34e-47e8-9e7d-3372be0b1691 · outbound

This paper cites DocVQA: A dataset for VQA on document images.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems DocVQA: A dataset for VQA on document images

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.656274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:38f502351e818809804337c8cb52ba42ddbf49c531d4186c0a2c1d4711ee7acf

Observation 3036f867-1409-4a0b-893f-5267038a278c · outbound

This paper cites InfographicVQA.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems InfographicVQA

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.646361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:2b34ffc33dad2bd3123f11faa1294cd2e307b3204f8ceda2a75e0fb367f8c7d7

Observation 1d88c7e6-7f08-4102-8ecc-3990a62686a1 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.647414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:e1ff1365d27b1fdad395b148a168340d769dc505dcf0a5d159a74bb994a3abd1

Observation 5cb459f8-d775-4233-b51c-dd4dcb55c463 · outbound

This paper cites Towards VQA models that can read.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Towards VQA models that can read

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.650510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:3d1c9e77cfb0d2f217fe2550a1cf5f8a77a928b85922499951bd2aec5295ceca

Observation fd5e739b-249b-42d0-9330-eeac67f2fecf · outbound

This paper cites UReader: Universal OCR-free visually situated language understanding with multimodal large language model.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems UReader: Universal OCR-free visually situated language understanding with multimodal large language model

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.652406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:a892ea573b3782acc189bc0900e887d80182afed95e83e9390dbb808e41bc614

Observation fd5df8c2-74ca-4768-a2d2-99b61bb009f7 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:13:06.652683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:75702c568eaf9782aa0f6382dc87143169aac71de09f9e48606f6ad6b3cb8405

Observation 680c603c-caa9-4fd0-9683-864d36dc43b4 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified structure learning for OCR-free document understanding.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems mPLUG-DocOwl 1.5: Unified structure learning for OCR-free document understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.624050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:f1e9b9cd9a1404ab5a6b80acc1059edc1897ea34e3950a2eaad4baeeb4f032a6

Observation ce0db268-efea-4acc-8ade-478b625b435d · outbound

This paper cites CogAgent: A visual language model for GUI agents.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems CogAgent: A visual language model for GUI agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.642170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:3f062ffdb6b7c63992f3bf10673dc846750e2b7d022294cc3172dfe4349e9864

Observation eeadb008-120f-4763-8649-c33e6465a655 · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Monkey: Image resolution and text label are important things for large multi-modal models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.654324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:28db7d33fa4d605c34cf7afb287312183de210aeb62a5300bf984e40ad84b042

Observation c1e3a305-7ffb-4e9a-9c0a-8ffd881c53d2 · outbound

This paper cites TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:13:06.655691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:a08aa3594c1bf16c8213598b04e9ebab0031d2b1696b14ad767ff948f7bbdfa6

Observation daf5c038-6a87-4bd2-ad6e-2659742359a8 · outbound

This paper cites TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:13:06.646603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:26a071d3097a390c040d5c2cf04515246a55812e0e6747a5bc0e389d12afdc5a

Observation 2d8c065b-c51d-464c-8043-09f285768c08 · outbound

This paper cites HRVDA: High-resolution visual document assistant.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems HRVDA: High-resolution visual document assistant

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.658689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:84e49ad9ef76b282d11bd3fc5f1075edbaeb2ab0bdaff50b787bcad64e163fb9

Observation b8946daf-1346-487b-ac91-46bc3862da70 · outbound

This paper cites Vary: Scaling up the vision vocabulary for large vision- language models.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Vary: Scaling up the vision vocabulary for large vision- language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.663478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:7adabdba5dd684e91ac57c8cbfde8764b764c6c6450f4a1c909e33673d5d63d1

Observation 771d640a-b0f3-432e-93bf-b2a82c800a1a · outbound

This paper cites MM1.5: Methods, analysis & insights from multimodal LLM fine-tuning.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems MM1.5: Methods, analysis & insights from multimodal LLM fine-tuning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.660592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:ab35b681f7edc0343cfda714398cc9a05dacb8ac68ecfe73aad943e9a883e7c5

Observation cdd41cf3-302b-4a0b-88df-8087a5e666be · outbound

This paper cites Marten: Visual question answering with mask generation for multi-modal document understanding.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Marten: Visual question answering with mask generation for multi-modal document understanding

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:13:23.642954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:b0605a3ff01e8203f54f23b017b1ea0695b2f329a9659f2bcb4b1645e6e9a485

Pith citing papers

Observation 4fea7622-0c1b-439b-9834-694c03244021 · inbound

When Correct Decisions Hide Internal Stress: Decision-State Probing in Multimodal Language Models cites this paper.

When Correct Decisions Hide Internal Stress: Decision-State Probing in Multimodal Language Models MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:17:26.216885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T19:02:12.613002Z digest=sha256:feb4f5d3eaca8c37834f3145095620504ef49bb12c80cf0058934a410ad4f9d2

Observation 324b48c5-fe9a-4e47-8523-173af80dcf8c · inbound

Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading cites this paper.

Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T04:25:12.551331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:25:12.551331Z digest=sha256:1ca7f3b60729d208325f4815d5cfb2c0cbb22dd5038ab6d960c960af75ffb0d5

Observation a8b007b0-eb69-4c28-b9b1-17c8f2156e70 · inbound

Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading cites this paper.

Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T04:34:36.466548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:34:36.466548Z digest=sha256:6bb4e31100c9cc322c8e56a889020fcd9f5bfd5382316004ccbb5da124c51d1e