Pith. sign in

Paper Citation Record · LEDGER

Describe Anything Model for Visual Question Answering on Text-rich Images

As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2507.12441.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12441 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:51:28.030074Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact2
  • verified fuzzy29
  • unresolved30
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84beee77-f1d4-4035-a9b7-5f7201e7181c · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Describe Anything Model for Visual Question Answering on Text-rich Images Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:22.229917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:22.229917Z digest=sha256:7dfa4cce95f1743c592241627542b06081a98458ae497067af4b2ccbfa5d35fb

Observation 79bac547-d27f-430e-944b-1006cabfc0d4 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Describe Anything Model for Visual Question Answering on Text-rich Images Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:22.309488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:22.309488Z digest=sha256:21ef8a162ca592077eb47b15a912e32dd1f6ab9f69ce6d87d7f33ef1915a1dba

Observation 7c81a785-24f2-415b-bfa1-4e5fd412585c · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Describe Anything Model for Visual Question Answering on Text-rich Images Lawrence Zitnick, and Devi Parikh

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:32.933048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:22.463704Z digest=sha256:2aaa77a86f0c52930fa9a19bac94357ed9d85db45d04cd925e42da7cd7240a2f

Observation e9cc3a14-c47a-477e-8694-381b28179564 · outbound

This paper cites Manmatha.

Describe Anything Model for Visual Question Answering on Text-rich Images Manmatha

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:32.796643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:22.554582Z digest=sha256:65b2f86b61bd7fd5d938c098c4d24745dac54112690b1730d51d0f5c972359dc

Observation 8a470535-b176-4f33-94f3-45f224946407 · outbound

This paper cites Manmatha.

Describe Anything Model for Visual Question Answering on Text-rich Images Manmatha

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:32.663645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:22.656721Z digest=sha256:5f70b4f6ea48579b9087b865d9f3cf3317457be8583bd865449d827272bba66d

Observation d8d2a888-5a6f-4ff0-9ad3-e569c7f39201 · outbound

This paper cites Qwen2.5-VL Technical Report.

Describe Anything Model for Visual Question Answering on Text-rich Images Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:22.733714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:22.733714Z digest=sha256:4bdd95e1fe00bd04a95446e0d727f07ea64eb88bee96c6e0fde2f18447e36d92

Observation 344dca3c-927f-4deb-86dc-f41c4a3ea717 · outbound

This paper cites Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation.

Describe Anything Model for Visual Question Answering on Text-rich Images Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:51:28.793646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:22.802127Z digest=sha256:c4c4b68028bd3a687a1f9db10cf4bf0f3f52d61d5dbd1b5ce9907520dccfc115

Observation 4ad8f5d5-e9bb-415d-b736-ca6e5a7cb673 · outbound

This paper cites Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee.

Describe Anything Model for Visual Question Answering on Text-rich Images Meyer, Yuning Chai, Dennis Park, and Yong Jae Lee

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:32.487993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:22.923107Z digest=sha256:85ff946459501d84493f64ba47defb0fe799d7a54ef4015476de92ff376f2dd5

Observation bf59f0d9-51d7-4de1-9dc9-850e09d2d67b · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Describe Anything Model for Visual Question Answering on Text-rich Images Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:22.997706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:22.997706Z digest=sha256:f5f6bda1fe4851be02058ae06658a467dc199cfca9a091956d9e060e4452f367

Observation 24af260d-1cfe-4dfd-a9b9-c4cffd26ca9c · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Describe Anything Model for Visual Question Answering on Text-rich Images InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:23.059337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:23.059337Z digest=sha256:c3c0c4064a7591f479434fc43121d93e380655368ac13f6633cf9b538824618c

Observation c2a536a4-87fd-4d88-82a1-083890949eaf · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Describe Anything Model for Visual Question Answering on Text-rich Images Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:23.141636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:23.141636Z digest=sha256:87608e616340b00ebb2ad0899a85a97c4ac5d78455aff3859ed3f32f89408026

Observation e726fda4-961a-49f9-9f94-11b812202f0d · outbound

This paper cites Scalable Vision Language Model Training via High Quality Data Curation.

Describe Anything Model for Visual Question Answering on Text-rich Images Scalable Vision Language Model Training via High Quality Data Curation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:23.196657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:23.196657Z digest=sha256:0117d9102071bcf89a83bff2f4b89d1b00d2b88309b64e3bcc537eac2d96f60f

Observation 1b8edb5d-19a6-4eb7-9eef-0f7e89541c57 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Describe Anything Model for Visual Question Answering on Text-rich Images Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:32.334006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:23.241332Z digest=sha256:16f06a25ab88a5daa68d2bf0438478c160dc52c53300491ee2e89683f9e573c1

Observation f5dd4f26-1fc6-476c-82ae-17cc33a5a2b1 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Describe Anything Model for Visual Question Answering on Text-rich Images A Survey on LLM-as-a-Judge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:23.278023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:23.278023Z digest=sha256:6f4bd81665aab1a30d156bfb1c96dc7056e6a0a139125765b1e3d15cca52f0d7

Observation 06b64956-6bb4-49af-9e10-b9e49bf46061 · outbound

This paper cites Regiongpt: Towards region understanding vision lan- guage model.

Describe Anything Model for Visual Question Answering on Text-rich Images Regiongpt: Towards region understanding vision lan- guage model

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:32.189072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:23.322295Z digest=sha256:e77fb5a1b1bb35a1366c2c13b4952385110564520c195c5919bd623a9eef84f3

Observation e8c18f00-3001-4fe4-886f-8417e463ac3e · outbound

This paper cites Hires-llava: Restoring fragmen- tation input in high-resolution large vision-language models.

Describe Anything Model for Visual Question Answering on Text-rich Images Hires-llava: Restoring fragmen- tation input in high-resolution large vision-language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:31.945422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:23.370201Z digest=sha256:47fb6cd89f753a31448d53fa416efffde6689fbc534d2e5e9370b279b7cf64e3

Observation 8449d0f0-9f88-4b7c-ba38-825c116fea0b · outbound

This paper cites Seeing out of the box: End- to-end pre-training for vision-language representation learn- ing.

Describe Anything Model for Visual Question Answering on Text-rich Images Seeing out of the box: End- to-end pre-training for vision-language representation learn- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:31.820105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:23.432894Z digest=sha256:b836cd8438873fbe9bc458fa1e91b0a7aa63be3926346d6d1ba58347400d5e68

Observation 2d1d844b-fbb9-4eba-9dd2-e89947ffb9d8 · outbound

This paper cites Multi-Agent VQA: Exploring Multi-Agent Foundation Models in Zero-Shot Visual Question Answering.

Describe Anything Model for Visual Question Answering on Text-rich Images Multi-Agent VQA: Exploring Multi-Agent Foundation Models in Zero-Shot Visual Question Answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:23.503836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:23.503836Z digest=sha256:7e2a3d41ca3314268a5742cb2bf1e8a3696b857bd328798fed2494be31b0bccb

Observation 00a07c69-fb59-494f-a0bb-a36d5b4123d9 · outbound

This paper cites Spa- tially aware multimodal transformers for textvqa.

Describe Anything Model for Visual Question Answering on Text-rich Images Spa- tially aware multimodal transformers for textvqa

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:31.696201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:23.575342Z digest=sha256:3be5d7c68dc692a1338c8be34b92626d8c5adc14e1416755920ab58e54991063

Observation a8f78605-3409-4d13-a3f9-d51b2fdcfc5f · outbound

This paper cites Binary codes capable of cor- recting deletions, insertions, and reversals.

Describe Anything Model for Visual Question Answering on Text-rich Images Binary codes capable of cor- recting deletions, insertions, and reversals

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:31.585020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:23.700616Z digest=sha256:9e5bf82aba81d90528756c6bf640b9355dd6fe47360c2599edf0cc75e72cb63b

Observation 92358a79-729d-47d1-aa8b-4cdc606fd0c0 · outbound

This paper cites From gener- ation to judgment: Opportunities and challenges of llm-as-a- judge.

Describe Anything Model for Visual Question Answering on Text-rich Images From gener- ation to judgment: Opportunities and challenges of llm-as-a- judge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:23.782830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:23.782830Z digest=sha256:e832a9f88cb16ec11ba4e28eaebf3e8163285a8e254b53151c9fad0fe68e2358

Observation 5ba75bd4-02e0-44dc-8bcf-3d33d8a96cdd · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Describe Anything Model for Visual Question Answering on Text-rich Images LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:23.854998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:23.854998Z digest=sha256:b95a88814cf9043a37dcb6db38976c533b86d3eb21796beabc0e9621c37ca56b

Observation 95cb28d3-56cf-49dc-9ab2-76a98c31f153 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Describe Anything Model for Visual Question Answering on Text-rich Images Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:23.950516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:23.950516Z digest=sha256:d8445da225f3f3e48b7346bead2645783804d69231cf94794ceef8bb8044cdfb

Observation 1b0da12a-8b7c-41f3-8183-55f5ebaf44b9 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with 9 frozen image encoders and large language models.

Describe Anything Model for Visual Question Answering on Text-rich Images Blip-2: Bootstrapping language-image pre-training with 9 frozen image encoders and large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:31.329965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:24.015831Z digest=sha256:17e3bcc2203511b8dae1750bfcbf25f9976c69ddf3c73570588327eb51e45f0f

Observation 733fa32e-9fdf-4b49-aa95-9a54c59959ea · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Describe Anything Model for Visual Question Answering on Text-rich Images VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:24.115929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:24.115929Z digest=sha256:79c9e4dfbc98f5ee5a44c3753370d6c94092ee743d5fda13c38aca4580c0f83c

Observation e4b0003d-c9ef-4039-aaf3-b683aa25f547 · outbound

This paper cites UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning.

Describe Anything Model for Visual Question Answering on Text-rich Images UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:24.195805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:24.195805Z digest=sha256:a38849f59091325810b7458abeb80eaf1635f40c13704d606f8031625e91c21e

Observation 7dedfffd-9ff3-42f3-95f0-e645bc29d63b · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

Describe Anything Model for Visual Question Answering on Text-rich Images Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:24.351412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:24.351412Z digest=sha256:257807969a73e4db5249701c51723302f94bb45a22e4f7bec2a65a4c112cf4eb

Observation 00944919-2683-4e18-9e2b-9d25779b7491 · outbound

This paper cites Enhancing visual document understanding with contrastive learning in large visual-language models.

Describe Anything Model for Visual Question Answering on Text-rich Images Enhancing visual document understanding with contrastive learning in large visual-language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:31.125927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:24.429859Z digest=sha256:26b01e736770c3823acde8e017cda133633f6c1a2d198eae20a15b63d835e0c8

Observation 786371db-77b2-4394-9012-dd66e803d6f6 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

Describe Anything Model for Visual Question Answering on Text-rich Images Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:30.961752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:24.520188Z digest=sha256:28134872741fb5476647f2ee5d680a9832c874746d85590624627b01754cd89a

Observation c5e311f2-055f-4a36-8c66-36da3105ce8f · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

Describe Anything Model for Visual Question Answering on Text-rich Images Describe Anything: Detailed Localized Image and Video Captioning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:24.655762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:24.655762Z digest=sha256:ba0864e72feeb73970bfbfaa86a36453aaa34687f5e04893556679e1e4423d1b

Observation 1309521c-a60a-444f-ba0b-13861ed211f6 · outbound

This paper cites Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos.

Describe Anything Model for Visual Question Answering on Text-rich Images Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:24.717294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:24.717294Z digest=sha256:a477cb8d6f0f5f284057794e0ee7a243ddd3e4499eb0f7427f9dfdfe352e8aeb

Observation 3b93f440-cb96-4770-bf11-64273b8b170f · outbound

This paper cites Revive: regional visual representa- tion matters in knowledge-based visual question answering.

Describe Anything Model for Visual Question Answering on Text-rich Images Revive: regional visual representa- tion matters in knowledge-based visual question answering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:30.758577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:24.785435Z digest=sha256:0f4355ca4c346e9276c60eceafb3c5825861c700b8cffcce0aca0cc8cb7ab54d

Observation b2244a65-2661-4e5a-9c3e-68e7cf52ac17 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Describe Anything Model for Visual Question Answering on Text-rich Images TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:25.051188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:25.051188Z digest=sha256:30bdd9a31f9dd4925eab4ae2bc7dc81898a040e62329e73cc843224f04346eb8

Observation a6538a19-2053-4f7b-9b23-425ec16e278e · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

Describe Anything Model for Visual Question Answering on Text-rich Images Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:25.447570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:25.447570Z digest=sha256:75a281c3208d4df49723dc26d8ea17ef0d90e014c4314d2d484c150418758f3b

Observation f4713053-4307-437d-9fb0-c1d5bf82f1c1 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Describe Anything Model for Visual Question Answering on Text-rich Images Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:25.555322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:25.555322Z digest=sha256:954d67b1d46a978be44c9d9df4f9a7d13266e5aab54ddfbfcbfb9bdc21445717

Observation e20bd04a-b6ce-4e0a-b61c-6265f5921053 · outbound

This paper cites Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning.

Describe Anything Model for Visual Question Answering on Text-rich Images Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:30.597058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:25.609867Z digest=sha256:d574677ec4573290726ed4ce6b25f49310ec5264fa021ef916ef4b200aaf0d00

Observation c027a98e-a9e1-4ea7-bf6c-cdc626865f15 · outbound

This paper cites ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering.

Describe Anything Model for Visual Question Answering on Text-rich Images ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:25.669672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:25.669672Z digest=sha256:834e52dbf7f506de6ecb8d61947ad21a4e703e7c43e4dcee20ffe6fbc8d523df

Observation 0ed54717-4d24-459a-8862-33932e98ed8f · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Describe Anything Model for Visual Question Answering on Text-rich Images Docvqa: A dataset for vqa on document images

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:30.453754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:25.772181Z digest=sha256:d844017d42244b06a1df539350773c0989b5e98fcd114127b0d0d4042b829bfa

Observation ed5aed7a-4b02-4b57-b8d6-d165feae8ca5 · outbound

This paper cites Infographicvqa.

Describe Anything Model for Visual Question Answering on Text-rich Images Infographicvqa

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:30.341084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:25.886738Z digest=sha256:c7e4c7f95e9f480026673370eeec4f9cf6e73036b915ecc63e2a7d92576da4cc

Observation 4bd88eda-5aae-4f02-9111-162a601db482 · outbound

This paper cites Im- proving automatic vqa evaluation using large language mod- els.

Describe Anything Model for Visual Question Answering on Text-rich Images Im- proving automatic vqa evaluation using large language mod- els

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:30.215862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:26.023556Z digest=sha256:2e15188d8f5d7798b97818ad4ceb5277096e3d8969a2dda83e87dc6ff7da241e

Observation aec207f3-7fb0-490b-9c7d-b048f9e70ac8 · outbound

This paper cites Dual dynamic consis- tency regularization for semi-supervised domain adaptation.

Describe Anything Model for Visual Question Answering on Text-rich Images Dual dynamic consis- tency regularization for semi-supervised domain adaptation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:30.060757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:26.157877Z digest=sha256:71cac51d0e01a2cae258d5fb03deb865cc18f413df5e98e35a7939a541bf498a

Observation 3f77f525-cd8e-49ef-8582-b112865b8d08 · outbound

This paper cites Enhancing Vietnamese VQA through Curriculum Learning on Raw and Augmented Text Representations.

Describe Anything Model for Visual Question Answering on Text-rich Images Enhancing Vietnamese VQA through Curriculum Learning on Raw and Augmented Text Representations

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:51:28.310404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:26.251341Z digest=sha256:39aa1411e029485cb23de2fe743e6e1e19b5aacf9db46bc341a30b70a10488b1

Observation 9e63eb81-1d27-45f8-905b-48b84c43d7ab · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Describe Anything Model for Visual Question Answering on Text-rich Images Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:26.400816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:26.400816Z digest=sha256:cfc1aaa84a3b6bb0ac9fc538d857926358d0ac61e6e95675e9cf0f34704414f1

Observation e0d03647-33cc-4a4f-be42-c6cb0f661b42 · outbound

This paper cites Going full-tilt boogie on document understanding with text-image-layout transformer.

Describe Anything Model for Visual Question Answering on Text-rich Images Going full-tilt boogie on document understanding with text-image-layout transformer

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:29.965392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:26.504317Z digest=sha256:c0967221fc4008fa1eab71c09da2994f2d75d59ae1acf9f1a3b152a28c51714f

Observation 90daa82d-0ea6-4b00-af00-80fb5d86b8e1 · outbound

This paper cites Towards vqa models that can read.

Describe Anything Model for Visual Question Answering on Text-rich Images Towards vqa models that can read

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:29.892501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:26.642120Z digest=sha256:8cb4bb883a934025536b7d6b3180e7da32ecef94ba3813f23d822fce8652c83a

Observation c66fd355-92af-4098-a040-fc45860ef684 · outbound

This paper cites Gemini 2.5: Pushing the frontier with ad- vanced reasoning, multimodality, long context, and next gen- eration agentic capabilities.

Describe Anything Model for Visual Question Answering on Text-rich Images Gemini 2.5: Pushing the frontier with ad- vanced reasoning, multimodality, long context, and next gen- eration agentic capabilities

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:29.825409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:26.791874Z digest=sha256:e236f6ac573659ec1af0f23030c93a401da60c4173df866e718fb49db8be5471

Observation d2fa1117-902e-4b5a-9b1f-80ce461457ff · outbound

This paper cites Igl-dt: Itera- tive global-local feature learning with dual-teacher semantic segmentation framework under limited annotation scheme.

Describe Anything Model for Visual Question Answering on Text-rich Images Igl-dt: Itera- tive global-local feature learning with dual-teacher semantic segmentation framework under limited annotation scheme

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:29.753418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:26.920416Z digest=sha256:7c32448055b09651f6cf162af1275f99968dc14edeb5cb849bf4df32b9d1abdb

Observation 245d17f9-558d-4cfc-b599-229a623e6672 · outbound

This paper cites Mlg2net: Molecular global graph network for drug response prediction in lung cancer cell lines.

Describe Anything Model for Visual Question Answering on Text-rich Images Mlg2net: Molecular global graph network for drug response prediction in lung cancer cell lines

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:29.657037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:27.048537Z digest=sha256:e0add82b05d9a35a4a8403df2e9ec73111e094e44a94b9b108917e98d2c924d6

Observation 4897f6b0-da76-4f82-b5a6-5728fd922136 · outbound

This paper cites Describe Anything in Medical Images.

Describe Anything Model for Visual Question Answering on Text-rich Images Describe Anything in Medical Images

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:27.152580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:27.152580Z digest=sha256:d6371e71013a708d6d7f896b0b5494e8cd4ed41b3f91e0c55b6a34824b40fc5d

Observation 3c94e104-6be6-41c1-ac8e-556f2f15b7ee · outbound

This paper cites Layoutlm: Pre-training of text and layout for document image understanding.

Describe Anything Model for Visual Question Answering on Text-rich Images Layoutlm: Pre-training of text and layout for document image understanding

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:29.556329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:27.234301Z digest=sha256:f3ae50e374f15bb26d4f24e7f7e433612d5bf356243aedf649fcf071985b067b

Observation eb9ae5ac-42fa-4da1-b241-8ea11a581587 · outbound

This paper cites LayoutLMv2: Multi-modal pre-training for visually-rich document under- standing.

Describe Anything Model for Visual Question Answering on Text-rich Images LayoutLMv2: Multi-modal pre-training for visually-rich document under- standing

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:29.460892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:27.346184Z digest=sha256:56daeed6920cca56e5a1ab4565c6815d987e475a661157ac7e3066a6ab8aee58

Observation b3db362a-53d8-482d-aefe-0a1772ec2630 · outbound

This paper cites Qwen3 Technical Report.

Describe Anything Model for Visual Question Answering on Text-rich Images Qwen3 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:27.484932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:27.484932Z digest=sha256:a303661c5f3f29f166d76b74e54da8a8bf59680809aa9e8790bcded932e1d5c5

Observation f42afdef-4402-4d9a-b516-53d06f90cedb · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Describe Anything Model for Visual Question Answering on Text-rich Images MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:27.556990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:27.556990Z digest=sha256:19f273bb7f2eef6ba37783a986200a154275ae2c5b2db3b591dd520c1af87d23

Observation a5e660d4-88a1-4d62-b350-74de87d435b9 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Describe Anything Model for Visual Question Answering on Text-rich Images Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:27.621969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:27.621969Z digest=sha256:b269d5cf34c7b8dcaeeea8412b3309b4c86f3c833e13479b37ff8b561038e43a

Observation 28d13954-ed74-4af8-affb-853296e5eb02 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Describe Anything Model for Visual Question Answering on Text-rich Images MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:27.706414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:27.706414Z digest=sha256:3ee3043934c7eae8564f16a59acfad0336bc84b0742e39e63122d0fe796a00c4

Observation 62c9845f-ed78-4449-bf61-0f9f206f4657 · outbound

This paper cites Osprey: Pixel un- derstanding with visual instruction tuning.

Describe Anything Model for Visual Question Answering on Text-rich Images Osprey: Pixel un- derstanding with visual instruction tuning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:29.320554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:27.807774Z digest=sha256:3ad2d808009a2aecf2d601412836ef4be6a5e59d226f678fdb58eb2162c4e0f2

Observation 88b25ccb-9f43-4a5b-9165-a65529b3b75f · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Describe Anything Model for Visual Question Answering on Text-rich Images VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:27.876456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:27.876456Z digest=sha256:d2bc56e7116d0027d101bc96a2390aa50de26c1f2bb76e70003fbba332d6d67b

Observation b847188f-6b8a-41bd-b0be-40ba08d13d9e · outbound

This paper cites Gpt4roi: Instruction tuning large language model on region- of-interest.

Describe Anything Model for Visual Question Answering on Text-rich Images Gpt4roi: Instruction tuning large language model on region- of-interest

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:29.166465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:27.936393Z digest=sha256:7960a0e769c00c3150d1f462cdc68240e5b7304dd71cb8c28550acf876dfeb30

Observation c5f36874-abf2-4723-9ad4-0522f873d8d7 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Describe Anything Model for Visual Question Answering on Text-rich Images Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:29.029064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T16:51:27.980688Z digest=sha256:ee6f2b1d3f27d6997e94370faf3efd39ef25efb1be50a59873902a10c8a58b23

Observation b30d1ff6-ecec-48bc-abdf-5041da93373e · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Describe Anything Model for Visual Question Answering on Text-rich Images InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:28.030074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:28.030074Z digest=sha256:69c39c70ecdddd46328fb55dfbc7ac97b22f4738afb0b807c3297171f5918a7e

Observation af5746fb-3597-4393-ac85-7f9f29d0aac8 · outbound

This paper cites an unresolved cited work.

Describe Anything Model for Visual Question Answering on Text-rich Images Unresolved cited work

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:27.417470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:27.417470Z digest=sha256:f136e3616290cf8b604165628a7a2ced48ef66178238dda50e92dda84436b451

Observation fec0588c-817c-4e8d-a524-d66f157feb45 · outbound

This paper cites an unresolved cited work.

Describe Anything Model for Visual Question Answering on Text-rich Images Unresolved cited work

Reference 2022

Resolution
parse uncertain
no resolver link, observed 2026-08-06T16:51:24.843669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:24.843669Z digest=sha256:890526d4eecd72daebc961f92205fe6436e6c7624fbedbd61a73fe38289d1d35

Pith citing papers

No inbound Pith citation observations are available.