Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T07:08:27.096029Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 3 inbound Pith citation observations for arXiv:2605.19307.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T07:08:27.096029Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T04:34:36.466548Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T22:17:26.215530Z
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 40a77528-40e2-4d85-a656-377706929da9 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems StoryLLaV A: Enhancing visual storytelling with multi-modal large language models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a32790f2-14ba-4cbe-a9c0-7e9a045cc975 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Refined semantic enhancement towards frequency diffusion for video captioning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e0a285ca-b2af-408a-9789-2ffe8a79dd31 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Action-aware linguistic skeleton optimization network for non-autoregressive video captioning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2cb57102-b457-4f1f-afab-e295cbfff7bb · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems VQA: Visual question answering
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 366adef9-3c3d-4ed0-9bff-125b5d6fcf40 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Making the V in VQA matter: Elevating the role of image understanding in visual question answering
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b8403d03-2906-4045-9077-0d8bc97bb874 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Robust visual question answering: Datasets, methods, and future challenges
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9b409ebb-1613-4226-8bee-c783736fb4ae · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Metamorphic Testing: A New Approach for Generating Next Test Cases
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a3cc4d09-97ab-47bc-bca5-3d667ce72d19 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems KVQA: Knowledge- aware visual question answering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7b58c1b2-cab4-4b3d-8450-7f94e5fc5db4 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems OCR-VQA: Visual question answering by reading text in images
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5a800ab6-245b-4278-a216-2196e09496fe · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Metamorphic testing: A review of challenges and opportunities
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation eba38e50-7dcd-4ad7-a71b-1bc60426b74b · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Perception matters: Detecting perception failures of VQA models using metamorphic testing
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b1ea812e-da81-412c-8d65-ec0f9b9bbedd · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Metamorphic testing of image captioning systems via image-level reduction
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation defa4c4a-2175-4611-ba56-e3c7800fc10c · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems How Multi-Modal LLMs Reshape Visual Deep Learning Testing? A Comprehensive Study Through the Lens of Image Mutation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b2b34d11-a9fa-41aa-b0ab-9409456710bf · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems CLIP in mirror: Disentangling text from visual images through reflection
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dc90bcf0-16f3-4835-91ad-7c23000b50c4 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9cad04b3-a962-45c0-b40b-1064c7646f41 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Improved baselines with visual instruction tuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d17fc3ed-8c8e-4869-ab75-3499af4d8a91 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7d250587-ca1f-4825-ba6a-0dc71fcb901f · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Qwen2.5-VL Technical Report
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f4dfff8c-3e48-48d0-9efa-f8f3188e200c · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3f4fcffa-d64c-4203-9598-e9edc0a06f77 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Cross-modal retrieval for knowledge-based visual question answering
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7299ff26-4b25-4e92-934d-911836cd0826 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems RoRA-VLM: Robust Retrieval-Augmented Vision Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e8c68954-cbd3-404f-b165-107084698c7f · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems EchoSight: Advancing visual-language models with wiki knowledge
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9863a52e-f48c-4a4f-be88-930f525f8f76 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Wiki-LLaV A: Hierarchical retrieval-augmented generation for multimodal LLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9d401e5e-077c-4c7f-b864-2c7217f1b398 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Augmenting multimodal LLMs with self-reflective tokens for knowledge- based visual question answering
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5259ee6c-ea88-4d77-949b-dbb3da6fd157 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Fine-Grained Knowledge Structuring and Retrieval for Visual Question Answering
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9537564a-1628-4075-be86-33cee8c52a11 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems MMKB-RAG: A Multi-Modal Knowledge-Based Retrieval-Augmented Generation Framework
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 13892cdb-6abd-4c60-a392-2b935ece5dc9 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Knowledge-based visual question answering with multimodal processing, retrieval, and filtering
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b93bf570-9e8d-4a8e-a89b-100242817fee · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Encyclopedic VQA: Visual questions about detailed properties of fine-grained categories
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 23acc7a2-4cb9-4009-8d5d-e0f358c54932 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Can pre-trained vision and language models answer visual information- seeking questions?
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a9766a17-b34e-47e8-9e7d-3372be0b1691 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems DocVQA: A dataset for VQA on document images
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3036f867-1409-4a0b-893f-5267038a278c · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems InfographicVQA
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1d88c7e6-7f08-4102-8ecc-3990a62686a1 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5cb459f8-d775-4233-b51c-dd4dcb55c463 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Towards VQA models that can read
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fd5e739b-249b-42d0-9330-eeac67f2fecf · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems UReader: Universal OCR-free visually situated language understanding with multimodal large language model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fd5df8c2-74ca-4768-a2d2-99b61bb009f7 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 680c603c-caa9-4fd0-9683-864d36dc43b4 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems mPLUG-DocOwl 1.5: Unified structure learning for OCR-free document understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ce0db268-efea-4acc-8ade-478b625b435d · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems CogAgent: A visual language model for GUI agents
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation eeadb008-120f-4763-8649-c33e6465a655 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Monkey: Image resolution and text label are important things for large multi-modal models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c1e3a305-7ffb-4e9a-9c0a-8ffd881c53d2 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation daf5c038-6a87-4bd2-ad6e-2659742359a8 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2d8c065b-c51d-464c-8043-09f285768c08 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems HRVDA: High-resolution visual document assistant
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b8946daf-1346-487b-ac91-46bc3862da70 · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Vary: Scaling up the vision vocabulary for large vision- language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 771d640a-b0f3-432e-93bf-b2a82c800a1a · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems MM1.5: Methods, analysis & insights from multimodal LLM fine-tuning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cdd41cf3-302b-4a0b-88df-8087a5e666be · outbound
MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems Marten: Visual question answering with mask generation for multi-modal document understanding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4fea7622-0c1b-439b-9834-694c03244021 · inbound
When Correct Decisions Hide Internal Stress: Decision-State Probing in Multimodal Language Models MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 324b48c5-fe9a-4e47-8523-173af80dcf8c · inbound
Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8b007b0-eb69-4c28-b9b1-17c8f2156e70 · inbound
Consistency Has a Computable Blind Spot: A Commutation Theory of Label-Free Reliability for Vision-Language Figure Reading MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.