Pith. sign in

Paper Citation Record · LEDGER

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code

As of 20 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 5 inbound Pith citation observations for arXiv:2412.02764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02764 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:11:05.881956Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:42:09.502274Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:18:32.707023Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1296daf9-43bf-443a-b070-0873db5177eb · outbound

This paper cites Le veraging large language models for data analysis automation,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code Le veraging large language models for data analysis automation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.237472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.676473Z digest=sha256:b2434288e44266ce80080fd2f14b74ade1a8d2c06feaa574576cf4a19f266bc6

Observation f805fbfb-e04f-4f06-aea0-0fb0557a9734 · outbound

This paper cites Pipe(line) dreams: Fully au tomated end- to-end analysis and visualization,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code Pipe(line) dreams: Fully au tomated end- to-end analysis and visualization,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.224123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.681412Z digest=sha256:2bfa054b8fd8ada85e27931efa3e6ab9359691d1504846d3f3a3c5435171f898

Observation 2641da3e-e307-485c-a8b1-6119d29ebe2a · outbound

This paper cites LLMs f or science: Usage for code generation and data analysis,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code LLMs f or science: Usage for code generation and data analysis,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.209946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.685741Z digest=sha256:58edeb540cf1ea8bdf62febae4d12cdf16f87a046fee9667e7e82820ece55c91

Observation f38cbc8b-4cae-4158-a8b9-8ea39c86bacf · outbound

This paper cites MatPlotAgent: Method and Evaluation for LLM-Based Agentic Scientific Data Visualization.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code MatPlotAgent: Method and Evaluation for LLM-Based Agentic Scientific Data Visualization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:05.689917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:05.689917Z digest=sha256:64baf1a7ab77bc481e05861c532b62b92a554ebaaeb9e13684aac332459eba18

Observation 8a23719f-2d9c-4cdc-a180-e040cfb01f1b · outbound

This paper cites Expectat ion vs. experi- ence: Evaluating the usability of code generation tools pow ered by large language models,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code Expectat ion vs. experi- ence: Evaluating the usability of code generation tools pow ered by large language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.196525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.694863Z digest=sha256:a047cfaf469f97cff8636033287f9d855ca1ad605b84f9b6384e6500f59755bf

Observation 80be52c5-7409-4286-8b65-f914ad241e1d · outbound

This paper cites DS-1000: A natural and reliable benchmark for data science code generation,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code DS-1000: A natural and reliable benchmark for data science code generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.182000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.699187Z digest=sha256:4c6d6d5ad6e9267b51f4e4ee61c0d3af96513a81b8df3276a1a9100ae3ccd954

Observation 88f2442b-7fdc-45b1-97c7-fe377454450e · outbound

This paper cites Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:05.703705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:05.703705Z digest=sha256:98d1c9c236b68626269508ed86555b9fe1e7fcce0d82c81a2b2aec3a397ac1e8

Observation 285ad928-ea9c-4a5a-8c1f-088447eb1363 · outbound

This paper cites ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:05.708253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:05.708253Z digest=sha256:9d9930080e28d0c2b406ad0a348f3183ce2322d6fb60b7e200238d0b7208a964

Observation 64cdd59e-fa45-4c7a-949e-f8a9b4c239ed · outbound

This paper cites nvBench: A Large-Scale Synthesized Dataset for Cross-Domain Natural Language to Visualization Task.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code nvBench: A Large-Scale Synthesized Dataset for Cross-Domain Natural Language to Visualization Task

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:05.712912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:05.712912Z digest=sha256:ac7c648c2b8f33503c48783cc659d742b1f9d3b160b9c480d4337601cdf1ffac

Observation ea5cffa7-83e3-413c-afd7-e957e792f133 · outbound

This paper cites pandas-dev/pandas: Pan das,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code pandas-dev/pandas: Pan das,

Reference 10

Resolution
verified exact
doi, observed 2026-08-11T23:11:05.916283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.717457Z digest=sha256:5219240ee5c9fe2af020bc06051e67aaca2dd63b4d6e895f667235851bca091f

Observation afea46eb-8104-4a19-923a-0d21c3029e9b · outbound

This paper cites Hello GPT-4o,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code Hello GPT-4o,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.168292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.721596Z digest=sha256:bca81f995d73e5f1ea132b61d4d4c2555d43f482d80f7c1a0a721de62e706a2b

Observation 0be27900-62e7-49e5-bbb4-8c82b56b9e1f · outbound

This paper cites The Claude 3 model family: Opus, Sonnet, Ha iku,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code The Claude 3 model family: Opus, Sonnet, Ha iku,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.153961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.725703Z digest=sha256:e23154e6737e662f1f5906616e097a70476bfbd6384b5d444dc0ac377963b2d5

Observation 0c463be8-7560-49c2-9990-bbd0599dfad4 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:05.834413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:05.834413Z digest=sha256:017060392c579f4321081159df33c015f6ccb4712cb1505240f1c02c4fca918e

Observation b3a9c9a4-9d28-44f1-8ea0-06a5c36d2422 · outbound

This paper cites The Llama 3 Herd of Models.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:05.840249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:05.840249Z digest=sha256:21fe878775e01988edd85bd3bb1debee90d97dd1988e50b005c46a12b1d1e815

Observation 16f84e66-a811-4db4-a582-6bfb983e1836 · outbound

This paper cites Matplotlib: A 2D graphics environment,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code Matplotlib: A 2D graphics environment,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.139185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.844633Z digest=sha256:27617924870b6811d7d275a70be7c8afe108468989bbd22e65bd1c7df8f1f54c

Observation dd5b4ecf-02e1-420b-9f32-e18687703228 · outbound

This paper cites Seaborn: statistical data visualizatio n,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code Seaborn: statistical data visualizatio n,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.126000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.848846Z digest=sha256:528e706e41bb0a61d24bd5fa64283acb514afc758a471dc3a69af700e455324f

Observation 42e82cb7-6a71-42f0-bc49-db8f253b1514 · outbound

This paper cites Collaborative data scienc e,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code Collaborative data scienc e,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.111677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.852967Z digest=sha256:7d6c8d0fe02bbe0d8289430e63d26b54edb1070df2bd7d4da4f0f1f99d834b57

Observation bce5986a-7059-4043-824a-3a4db4378bd4 · outbound

This paper cites HuggingFace Page with PandasPlotBench,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code HuggingFace Page with PandasPlotBench,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.096229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.856987Z digest=sha256:c1c357b323dd754172ddd20fa4ff44aadd327a99c93bd4bea19b075aa42957cd

Observation 03bdaba8-08a6-4907-9d4d-cf77ba6c765a · outbound

This paper cites Code for running the benchmark,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code Code for running the benchmark,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.081942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.861210Z digest=sha256:af82ec44e4fa474d53ebf34a68377c49e9696966eee56d49883c46f7ab73dd39

Observation 98e80da5-cb13-4b0a-9fa8-95c1c474fb69 · outbound

This paper cites Supplementary materials,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code Supplementary materials,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.067355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.865295Z digest=sha256:2d037ae23606067ab709cf83a6ac84e5ae978ada1150189050832f33f4736557

Observation 4c155640-331a-40c7-90d4-91840aaf6440 · outbound

This paper cites MatPlotLib Gallery,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code MatPlotLib Gallery,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.053973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.869621Z digest=sha256:6b263db59bfb38a11647002317b0f1f64efa4d31f78dd379bce4c1149e4f44a8

Observation bb151c14-0d12-4129-9fb0-3efcb53b472f · outbound

This paper cites GPT-4 Technical Report.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code GPT-4 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:05.873704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:05.873704Z digest=sha256:c36f07e76b00274f6d5c8af4d46b281b6b05fccea34d78de83024ae68d24854f

Observation 2ad7261f-4ab2-436b-8bf1-f5bfb3363170 · outbound

This paper cites GPT-4V(ision) System Card,.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code GPT-4V(ision) System Card,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:11:06.039487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T23:11:05.877858Z digest=sha256:402d9e9f6a9dedca0701cc1c2e4eb0567df98d54d19da9f517489030e6fa3484

Observation 9188646f-6a84-4200-99b0-a6bd27450897 · outbound

This paper cites CodeBERT: A Pre-Trained Model for Programming and Natural Languages.

Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code CodeBERT: A Pre-Trained Model for Programming and Natural Languages

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T23:11:05.881956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:11:05.881956Z digest=sha256:74519650431706636b3ca7ec224da6e520b81bba900d6f1d74dd228780638ad2

Pith citing papers

Observation 44978140-8e11-45f5-8954-b165103bdd0c · inbound

MLDebugging: Towards Benchmarking Code Debugging Across Multi-Library Scenarios cites this paper.

MLDebugging: Towards Benchmarking Code Debugging Across Multi-Library Scenarios Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:09.502274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:09.502274Z digest=sha256:1249509af14882f042f981aaa5673040af98d8bba433ac15e8a1bf0084eed706

Observation 62edd80c-aa5f-41dc-ad03-6e0a71601121 · inbound

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction cites this paper.

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:11:27.539239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T11:06:38.102348Z digest=sha256:ed6615bbc9f71e9ec7b12cbd239cf0d2e7b061b792fa1024e5d0f9c66ad94008

Observation cdfda063-5c8d-4404-8dbb-bc397fcba74a · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T15:55:53.399860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:55:53.399860Z digest=sha256:136a8f473e1a545ca4fbe3cd68a5af08a6d971ea9e9ade39efa6a0c20960055b

Observation 7967d04d-d35c-4235-892b-d00caa6a0a85 · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T17:09:17.931502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:09:17.931502Z digest=sha256:d191c54a29cc229bccafa8d7ef843cf0611eb43e5e2e555b4427ed35740c92c3

Observation 2075e0c2-d607-4df1-94aa-2bc1814b3fcb · inbound

PairCoder++: Pair Programming as a Universal Paradigm for Verified Code-Driven Multimodal and Structured-Artifact Generation cites this paper.

PairCoder++: Pair Programming as a Universal Paradigm for Verified Code-Driven Multimodal and Structured-Artifact Generation Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T15:18:32.709074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-03T15:15:14.024594Z digest=sha256:5cca087a20c477573df6d9aa143c5ce4a7805190bace75d5ea3b13d80c4afe28