Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:51:28.030074Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2507.12441.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:51:28.030074Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 84beee77-f1d4-4035-a9b7-5f7201e7181c · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79bac547-d27f-430e-944b-1006cabfc0d4 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c81a785-24f2-415b-bfa1-4e5fd412585c · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Lawrence Zitnick, and Devi Parikh
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e9cc3a14-c47a-477e-8694-381b28179564 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Manmatha
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a470535-b176-4f33-94f3-45f224946407 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Manmatha
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d8d2a888-5a6f-4ff0-9ad3-e569c7f39201 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Qwen2.5-VL Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 344dca3c-927f-4deb-86dc-f41c4a3ea717 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ad8f5d5-e9bb-415d-b736-ca6e5a7cb673 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bf59f0d9-51d7-4de1-9dc9-850e09d2d67b · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24af260d-1cfe-4dfd-a9b9-c4cffd26ca9c · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2a536a4-87fd-4d88-82a1-083890949eaf · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e726fda4-961a-49f9-9f94-11b812202f0d · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Scalable Vision Language Model Training via High Quality Data Curation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b8edb5d-19a6-4eb7-9eef-0f7e89541c57 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f5dd4f26-1fc6-476c-82ae-17cc33a5a2b1 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images A Survey on LLM-as-a-Judge
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06b64956-6bb4-49af-9e10-b9e49bf46061 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Regiongpt: Towards region understanding vision lan- guage model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e8c18f00-3001-4fe4-886f-8417e463ac3e · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Hires-llava: Restoring fragmen- tation input in high-resolution large vision-language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8449d0f0-9f88-4b7c-ba38-825c116fea0b · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Seeing out of the box: End- to-end pre-training for vision-language representation learn- ing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2d1d844b-fbb9-4eba-9dd2-e89947ffb9d8 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Multi-Agent VQA: Exploring Multi-Agent Foundation Models in Zero-Shot Visual Question Answering
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00a07c69-fb59-494f-a0bb-a36d5b4123d9 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Spa- tially aware multimodal transformers for textvqa
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a8f78605-3409-4d13-a3f9-d51b2fdcfc5f · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Binary codes capable of cor- recting deletions, insertions, and reversals
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 92358a79-729d-47d1-aa8b-4cdc606fd0c0 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images From gener- ation to judgment: Opportunities and challenges of llm-as-a- judge
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ba75bd4-02e0-44dc-8bcf-3d33d8a96cdd · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95cb28d3-56cf-49dc-9ab2-76a98c31f153 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b0da12a-8b7c-41f3-8183-55f5ebaf44b9 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Blip-2: Bootstrapping language-image pre-training with 9 frozen image encoders and large language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 733fa32e-9fdf-4b49-aa95-9a54c59959ea · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images VisualBERT: A Simple and Performant Baseline for Vision and Language
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4b0003d-c9ef-4039-aaf3-b683aa25f547 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dedfffd-9ff3-42f3-95f0-e645bc29d63b · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Oscar: Object-semantics aligned pre-training for vision-language tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00944919-2683-4e18-9e2b-9d25779b7491 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Enhancing visual document understanding with contrastive learning in large visual-language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 786371db-77b2-4394-9012-dd66e803d6f6 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Mon- key: Image resolution and text label are important things for large multi-modal models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c5e311f2-055f-4a36-8c66-36da3105ce8f · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Describe Anything: Detailed Localized Image and Video Captioning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1309521c-a60a-444f-ba0b-13861ed211f6 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b93f440-cb96-4770-bf11-64273b8b170f · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Revive: regional visual representa- tion matters in knowledge-based visual question answering
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b2244a65-2661-4e5a-9c3e-68e7cf52ac17 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6538a19-2053-4f7b-9b23-425ec16e278e · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4713053-4307-437d-9fb0-c1d5bf82f1c1 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Ovis: Structural Embedding Alignment for Multimodal Large Language Model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e20bd04a-b6ce-4e0a-b61c-6265f5921053 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c027a98e-a9e1-4ea7-bf6c-cdc626865f15 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ed54717-4d24-459a-8862-33932e98ed8f · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Docvqa: A dataset for vqa on document images
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ed5aed7a-4b02-4b57-b8d6-d165feae8ca5 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Infographicvqa
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4bd88eda-5aae-4f02-9111-162a601db482 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Im- proving automatic vqa evaluation using large language mod- els
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aec207f3-7fb0-490b-9c7d-b048f9e70ac8 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Dual dynamic consis- tency regularization for semi-supervised domain adaptation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3f77f525-cd8e-49ef-8582-b112865b8d08 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Enhancing Vietnamese VQA through Curriculum Learning on Raw and Augmented Text Representations
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9e63eb81-1d27-45f8-905b-48b84c43d7ab · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0d03647-33cc-4a4f-be42-c6cb0f661b42 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Going full-tilt boogie on document understanding with text-image-layout transformer
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 90daa82d-0ea6-4b00-af00-80fb5d86b8e1 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Towards vqa models that can read
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c66fd355-92af-4098-a040-fc45860ef684 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Gemini 2.5: Pushing the frontier with ad- vanced reasoning, multimodality, long context, and next gen- eration agentic capabilities
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d2fa1117-902e-4b5a-9b1f-80ce461457ff · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Igl-dt: Itera- tive global-local feature learning with dual-teacher semantic segmentation framework under limited annotation scheme
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 245d17f9-558d-4cfc-b599-229a623e6672 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Mlg2net: Molecular global graph network for drug response prediction in lung cancer cell lines
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4897f6b0-da76-4f82-b5a6-5728fd922136 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Describe Anything in Medical Images
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c94e104-6be6-41c1-ac8e-556f2f15b7ee · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Layoutlm: Pre-training of text and layout for document image understanding
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eb9ae5ac-42fa-4da1-b241-8ea11a581587 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images LayoutLMv2: Multi-modal pre-training for visually-rich document under- standing
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b3db362a-53d8-482d-aefe-0a1772ec2630 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Qwen3 Technical Report
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f42afdef-4402-4d9a-b516-53d06f90cedb · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5e660d4-88a1-4d62-b350-74de87d435b9 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28d13954-ed74-4af8-affb-853296e5eb02 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62c9845f-ed78-4449-bf61-0f9f206f4657 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Osprey: Pixel un- derstanding with visual instruction tuning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 88b25ccb-9f43-4a5b-9165-a65529b3b75f · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b847188f-6b8a-41bd-b0be-40ba08d13d9e · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Gpt4roi: Instruction tuning large language model on region- of-interest
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c5f36874-abf2-4723-9ad4-0522f873d8d7 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b30d1ff6-ecec-48bc-abdf-5041da93373e · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af5746fb-3597-4393-ac85-7f9f29d0aac8 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Unresolved cited work
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fec0588c-817c-4e8d-a524-d66f157feb45 · outbound
Describe Anything Model for Visual Question Answering on Text-rich Images Unresolved cited work
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.