Pith. sign in

Paper Citation Record · LEDGER

AIDE: Agentically Improve Visual Language Model with Domain Experts

As of 8 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2502.09051.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09051 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T22:51:02.176089Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:33:56.525727Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:34:04.453845Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9fd34b14-1a2d-4600-be48-0b6c5d838139 · outbound

This paper cites https://github.com/PaddlePaddle/PaddleOCR.

AIDE: Agentically Improve Visual Language Model with Domain Experts https://github.com/PaddlePaddle/PaddleOCR

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T22:51:02.599280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T22:51:02.029144Z digest=sha256:55082318234a46668ea0b1b89e7458ded2e914feb49f8a0b1f03a1446f728824

Observation e3a99080-5ac5-41f9-a351-04e6bdf735ed · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

AIDE: Agentically Improve Visual Language Model with Domain Experts Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.035005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.035005Z digest=sha256:632258c8dce6b42220452cb682215f2cd554a69bc76fa3ee3e367508edf88c4d

Observation 477853ad-ab42-431a-8877-4c30b5d58839 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

AIDE: Agentically Improve Visual Language Model with Domain Experts ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.040957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.040957Z digest=sha256:b23603ac2cbf56517e5ce21ba32a72df7c1b48bd79eef3fe273c8be82c5f7383

Observation ae84e11c-25a2-49fc-8504-0c76cd8e3e6b · outbound

This paper cites ColorSense: A Study on Color Vision in Machine Visual Recognition.

AIDE: Agentically Improve Visual Language Model with Domain Experts ColorSense: A Study on Color Vision in Machine Visual Recognition

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T22:51:02.454498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T22:51:02.047493Z digest=sha256:2534572a755c183cd1d655e5c61c4ca216a44a9107a05b54b99a7506c6a43326

Observation 288263e2-4415-4cd8-973d-2d77755e6549 · outbound

This paper cites MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models.

AIDE: Agentically Improve Visual Language Model with Domain Experts MegaCOIN: Enhancing Medium-Grained Color Perception for Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.054133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.054133Z digest=sha256:7f47cd1d0bbd0996a58be894377f8d1381dac7dc38cc54ae7ee9f1f4a381e983

Observation 707eaddc-99b1-4369-a018-e021d1ce2148 · outbound

This paper cites VILA$^2$: VILA Augmented VILA.

AIDE: Agentically Improve Visual Language Model with Domain Experts VILA$^2$: VILA Augmented VILA

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.060115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.060115Z digest=sha256:1e1d5dc5d1252f152356ffef04b58cce5ebcdef5a12168148561966d5f38fe51

Observation 17b99959-2ea4-4a07-b09e-aacff4629978 · outbound

This paper cites an unresolved cited work.

AIDE: Agentically Improve Visual Language Model with Domain Experts Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T22:51:02.582399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T22:51:02.066728Z digest=sha256:dc7a0ef17ab8e2df8003f4a3b725e5deb7190e585a2cd3a377d0380a75b35862

Observation 86d13882-21a5-4091-8838-69c0b071b24b · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

AIDE: Agentically Improve Visual Language Model with Domain Experts MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.077942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.077942Z digest=sha256:474615d1acc1e7e52c9171de3f95ddfde1dd2716ed8e851584e5e2ac237d44d2

Observation 539149fd-9897-41cf-b11c-8f42aeac7118 · outbound

This paper cites Language Is Not All You Need: Aligning Perception with Language Models.

AIDE: Agentically Improve Visual Language Model with Domain Experts Language Is Not All You Need: Aligning Perception with Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.083431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.083431Z digest=sha256:66d3a30e40cde575b4f461b449433b60a726e19d760dedbd5e64be1987a45f50

Observation 54530509-9881-4191-8e03-df9c6846c256 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

AIDE: Agentically Improve Visual Language Model with Domain Experts Evaluating Object Hallucination in Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.089154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.089154Z digest=sha256:5a85f36647ddfe2669f522351aa70bc3bee9d0f256d82766097545051bb0b45b

Observation 8bf98f31-0817-465a-842c-6028598edeba · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

AIDE: Agentically Improve Visual Language Model with Domain Experts Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.094840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.094840Z digest=sha256:ca27bc3e635b743eaa58718bd830900a9bbcd2970c57499c92f4cda4ec2bfe85

Observation 6d1c5df0-0697-4f7b-bf31-2729069bfb5f · outbound

This paper cites MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning.

AIDE: Agentically Improve Visual Language Model with Domain Experts MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.100050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.100050Z digest=sha256:7f146175c60b1057c47e19030af6235bf2338a9cc9d24077d54341dff1858fc4

Observation 32e65011-305e-4243-94c0-f70778bc5d39 · outbound

This paper cites Visual Instruction Tuning.

AIDE: Agentically Improve Visual Language Model with Domain Experts Visual Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.105296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.105296Z digest=sha256:f4b72970faef99b2c22c4245c64b08c18175d8d1be3db190b059d3b593ded652

Observation 801101c5-25ef-4aa6-bb36-b682c8af1162 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

AIDE: Agentically Improve Visual Language Model with Domain Experts Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.110594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.110594Z digest=sha256:985e46c67e3cec47ec14a5d52b7810bca9f13ca68678c0dd74d3c835e38acb4b

Observation a95793cf-c7a3-48be-930d-3fc39b394dab · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

AIDE: Agentically Improve Visual Language Model with Domain Experts MMBench: Is Your Multi-modal Model an All-around Player?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.115887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.115887Z digest=sha256:c4b6f81b7772554b5d0af6bfd334d6b74d524f59c29c48a8093c3ca793751fc8

Observation 35d0c92a-b5ab-48f1-bf2e-0f1d10152ea7 · outbound

This paper cites an unresolved cited work.

AIDE: Agentically Improve Visual Language Model with Domain Experts Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T22:51:02.566521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T22:51:02.120928Z digest=sha256:e300526c29c7528271f13c5625854c9bfc925814b96e10098e9ad422ea903086

Observation a3334bfa-f0a5-46b5-89b3-b804ec73ca3b · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

AIDE: Agentically Improve Visual Language Model with Domain Experts ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.125610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.125610Z digest=sha256:f466bb1e1c7cbb6ac46d98137f9a2d55302d274f70f5f2c4e08c592cac76fb00

Observation 7b7deb53-b14d-415b-885b-03ae6c5fcc8f · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

AIDE: Agentically Improve Visual Language Model with Domain Experts Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.130476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.130476Z digest=sha256:ce40882105a84b808fefc435b1842b886a1683b893ee9b1eee97d9d7be57fe92

Observation c07c2caa-faa4-46f9-ae1d-b4c4accee154 · outbound

This paper cites an unresolved cited work.

AIDE: Agentically Improve Visual Language Model with Domain Experts Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T22:51:02.551802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T22:51:02.135605Z digest=sha256:9d25068f6c917671119cc5b9578b16ee86694f624e618f712a5a19b99e2d27e5

Observation 9e50171f-9fac-4c3c-a3d1-e69360510242 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

AIDE: Agentically Improve Visual Language Model with Domain Experts Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.140407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.140407Z digest=sha256:5c53a341f39dfaea81a6ac7ee236bc892a912059f5e1e03e969d2b7dc94c8d87

Observation 29b874c9-f940-4a09-b68b-1e0d9b484a78 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

AIDE: Agentically Improve Visual Language Model with Domain Experts Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.145279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.145279Z digest=sha256:7bdf008907b357e6d1a32541c547a709cdf979a606b13134a9f2af5c5a8c7107

Observation 9288146e-b192-4e30-99dc-fb5f06424d32 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

AIDE: Agentically Improve Visual Language Model with Domain Experts Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.150470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.150470Z digest=sha256:3dc0b4992f568f7631539c48ea897cab027acc3db9631bb931bbecd6b4597170

Observation 425c3e3c-6699-4e6a-b98f-4afc0b156f14 · outbound

This paper cites an unresolved cited work.

AIDE: Agentically Improve Visual Language Model with Domain Experts Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T22:51:02.536797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T22:51:02.155520Z digest=sha256:46a0086801e1c53ee586e256954e326f49f6f79000bf4781416463adbdae1811

Observation b37e27cd-d0da-491f-9f3c-1e4c4349c0ff · outbound

This paper cites an unresolved cited work.

AIDE: Agentically Improve Visual Language Model with Domain Experts Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T22:51:02.521494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T22:51:02.160317Z digest=sha256:4f44e41b49f2ce45ea8b9614bfbc78af26381854701588ee0b637853e9322f69

Observation c062180a-2661-4a6b-a21e-c281cf7ac7ec · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

AIDE: Agentically Improve Visual Language Model with Domain Experts Florence: A New Foundation Model for Computer Vision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.164934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.164934Z digest=sha256:f6050a8a8963358935af835b0f20318595b13fb2f34a76be247696061ec1b231

Observation 6c9d1c0e-3121-461b-83d5-51f24c97b64d · outbound

This paper cites online" 'onlinestring :=.

AIDE: Agentically Improve Visual Language Model with Domain Experts online" 'onlinestring :=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.170241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.170241Z digest=sha256:ea8593120e25d35eef72189bb20d821260bbc64945a85fb3cef1613ad70dcdbd

Observation 76938200-1074-44f6-b82a-50bf4817e6ed · outbound

This paper cites write newline.

AIDE: Agentically Improve Visual Language Model with Domain Experts write newline

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T22:51:02.176089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:51:02.176089Z digest=sha256:98231a67f7d0f8d1bf72cd1ea8e458dd5cca2d182cfdfeb99726608aa61c8a0d

Pith citing papers

Observation 7d51a540-97a0-41d0-9202-02ef0d0b3ea4 · inbound

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training cites this paper.

VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training AIDE: Agentically Improve Visual Language Model with Domain Experts

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:34:04.544547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:33:56.525727Z digest=sha256:d26575b3280ea01c56fe602390c0e5bac3bb4c863e8accdbc23e158333a00fc8