Pith. sign in

Paper Citation Record · LEDGER

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering

As of 7 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2510.04514.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.04514 v3

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:32:16.446676Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 60ef1053-64f9-478a-996a-ed69c7de5f2d · outbound

This paper cites • Normalize both answers carefully by checking both the <groundtruth answer> and <predicted answer>values in context.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering • Normalize both answers carefully by checking both the <groundtruth answer> and <predicted answer>values in context

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:16.076536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:16.076536Z digest=sha256:ff180c0f6911764a1ce15e6781d5bf708c0dc7b9f58b694ad642f4864263d3e9

Observation fe88b479-d16b-4cc0-aa09-2ef8d1fc4873 · outbound

This paper cites ground_truth_filtered.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering ground_truth_filtered

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:16.197258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:16.197258Z digest=sha256:3dc52ca25ca2436dff2d7f3df094147973c7a81b42c5f3fc53d69b7ea9299c18

Observation 2d2c7b92-dfaf-4ef0-8d19-ae06ade311f6 · outbound

This paper cites How much higher is X than Y?.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering How much higher is X than Y?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:16.446676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:16.446676Z digest=sha256:a98f0bdc479710971f22ddcf6c7bbea2357512ad28dcadb2a70f9a33bc26f488

Observation d7c817ed-9b54-469a-930d-4fa99cd3cb0f · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:14.675403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:14.675403Z digest=sha256:96584c323199d65d741fa01895b8f902194709c1f9f16b3fba9f86c039e637b8

Observation fdceb79e-7fb5-4a9b-9560-81c34931c085 · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering SmolVLM: Redefining small and efficient multimodal models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:14.892848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:14.892848Z digest=sha256:79a7b592ff6ce62db1f5d290bd008e67760977378c45470b0f43e5a4f3e0502f

Observation fcee366a-7990-4787-8542-9bed575f0004 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:14.983970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:14.983970Z digest=sha256:f1be614953062e3254ce2bcf3c4cd6fc2277a3224dbdc2cc9f49570e75fb10de

Observation 07be4d4d-cc99-4fee-9539-94cc41f6ed51 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:15.101569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:15.101569Z digest=sha256:3b826171d28bf4f6b95b4eb2c6c0367db9e06201f98b9a9b2629b6636d1effa8

Observation a095ebf5-ebae-4d08-af40-75a0eb04acbe · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:15.178796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:15.178796Z digest=sha256:693afcea5f6bdf3b2c24d5eca3a6722f939669bf3b8bcf19560476606fdd50b7

Observation 5d46db7c-a6ee-4af3-9dd8-e3c451bd7ffb · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:15.292166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:15.292166Z digest=sha256:8624ce5d55703d46821377b32fbde186cbdabde2974c7742775e62a91d8e58d4

Observation e9cef323-f9b7-46b6-b53b-a8834caf2566 · outbound

This paper cites What is the value of India in 2021?.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering What is the value of India in 2021?

Reference 11

Resolution
malformed identifier
no resolver link, observed 2026-08-04T11:32:15.405966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:15.405966Z digest=sha256:31bbf604f0b2d8a1173c3cf0f5168d3efd6c7f6975e98e5ca2b813e272b4c7fd

Observation 48ab2163-39e2-4acb-8237-c2e57d82b855 · outbound

This paper cites in thousands.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering in thousands

Reference 16

Resolution
malformed identifier
no resolver link, observed 2026-08-04T11:32:15.923468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:15.923468Z digest=sha256:09162c2833882b2517551aacb5b7db2d7c4d26182120e8bb38902ce36dd2a647

Observation 0421bbad-3ea9-4f79-a6db-165245af6cf1 · outbound

This paper cites an unresolved cited work.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:16.279348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:16.279348Z digest=sha256:00018542c30c52669670d9a1b28144054592ecbec3150a0bc301720299a74da3

Observation dad9b0ea-790e-4c3f-9fe9-aa35b7fa9af1 · outbound

This paper cites an unresolved cited work.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:16.377453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:16.377453Z digest=sha256:f95e2fb733b5a87c3df2f284fe197d868205f48e821e5cd83d79bc1becf55b6a

Observation 97be4076-f69b-424e-9617-5991f95549aa · outbound

This paper cites For segmentation tasks, we use the Segment Anything (ViT-H) (Kirillov et al.,.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering For segmentation tasks, we use the Segment Anything (ViT-H) (Kirillov et al.,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:15.794019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:15.794019Z digest=sha256:0fe8d4e6c009cccbe42af65a4fe0e3410a7e4e913beec58b9f379000dd9e1d0d

Observation 3b86e36d-fb5c-41d9-9acd-983ca0ce1de1 · outbound

This paper cites (b) Chart-specific Tools Localized the X-axis and applied pixel-to-value interpolation to compute the median from the left and right parts of the detected box of interest.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering (b) Chart-specific Tools Localized the X-axis and applied pixel-to-value interpolation to compute the median from the left and right parts of the detected box of interest

Reference 203

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:15.614374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:15.614374Z digest=sha256:29de3beeb3d5a70179bdf79b7c43ee386e8e6e451373f62cb048d1d394d1af73

Observation bbd06f16-89e3-4c17-9bf5-751bb1300e9e · outbound

This paper cites The tools that utilize EasyOCR are the same as above.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering The tools that utilize EasyOCR are the same as above

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:15.480923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:15.480923Z digest=sha256:1689d1e39b38aadfe5527a144ad2af2b1db921338343ab2a010214e28de91803

Observation a96229d3-480b-4d31-9e3c-eafbbaf8dd64 · outbound

This paper cites Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:14.377050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:14.377050Z digest=sha256:b63f317772bf564a9a33fb6d698a97d2af5a5c54366d41ee0511bfa7eba545f5

Observation 332aad73-b969-46c7-9c8e-e681d7075fdb · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:14.531541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:14.531541Z digest=sha256:a57b2d4f24a132003b81207afbb7d973da8a477e21ace3b168e8d885fa9c7587

Observation e51e3b83-539a-4c3f-9990-228ba536ca0d · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:14.317722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:14.317722Z digest=sha256:f68ed8af52f40104340aebec8034ddd3e4eedd687188a53c49e978fff33c141d

Observation 7dd154c6-d455-4066-8564-f221c43666c4 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:14.796214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:14.796214Z digest=sha256:1b82384b2ec366bd2e09cfcd7373d80acbbd7f30675dc5cdd5efa10a383b78e2

Observation 92625edb-977d-4237-b0a9-dd6fee19e61e · outbound

This paper cites an unresolved cited work.

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering Unresolved cited work

Reference 2409

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:15.690901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:15.690901Z digest=sha256:367616aa5c7e231bd0860c65e29e52817aa7f7b6d54850fc566964db3844c4f4

Pith citing papers

No inbound Pith citation observations are available.