Pith. sign in

Paper Citation Record · LEDGER

A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2305.18486.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.18486 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:21:25.996298Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

13
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 86fe6ce0-57e2-43b6-9055-5afff4a46d8a · inbound

A Survey on Retrieval-Augmented Text Generation for Large Language Models cites this paper.

A Survey on Retrieval-Augmented Text Generation for Large Language Models A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:15:55.252575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T02:15:05.379583Z digest=sha256:30a53ad81eefb68f553daf8c50fb8d067bc89bc6272831a9d38186861f79c75c

Observation ee5747c1-46fa-4e75-8990-c261c810452c · inbound

An Empirical Study of Evaluating Long-form Question Answering cites this paper.

An Empirical Study of Evaluating Long-form Question Answering A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:21:25.996298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:21:25.996298Z digest=sha256:fa67aa5d3bead60deaba296b506f0e8d308ecbefe6e8313396340adcec1d9d0c

Observation 53970024-7059-452c-8efb-5401dbe97f31 · inbound

Fine-Tuning Large Language Models and Evaluating Retrieval Methods for Improved Question Answering on Building Codes cites this paper.

Fine-Tuning Large Language Models and Evaluating Retrieval Methods for Improved Question Answering on Building Codes A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:41:18.234345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:41:18.234345Z digest=sha256:7817e13ded2aa5ef270a4a39063105e4bbbacb7cdd10fda08c61a8a98b0d89f1

Observation c2b25631-689f-4039-a697-f7538705028e · inbound

Evaluation of LLMs for mathematical problem solving cites this paper.

Evaluation of LLMs for mathematical problem solving A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:06.862023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:06.862023Z digest=sha256:a64526159a6832b751a516978e8ace337b4b756a9c2ac87d7de2748c260fe675

Observation a32a603d-58f8-424e-892d-0bdd52441206 · inbound

Intent Matters: Enhancing AI Tutoring with Fine-Grained Pedagogical Intent Annotation cites this paper.

Intent Matters: Enhancing AI Tutoring with Fine-Grained Pedagogical Intent Annotation A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:33:26.943474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:33:26.943474Z digest=sha256:dee665a4b10abf80391141efb98b40fc2499f67e20630cdb62599c2954e56b15

Observation 26886995-d1a5-40bb-b637-57ade7dbad5f · inbound

Fusing Knowledge and Language: A Comparative Study of Knowledge Graph-Based Question Answering with LLMs cites this paper.

Fusing Knowledge and Language: A Comparative Study of Knowledge Graph-Based Question Answering with LLMs A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T19:24:53.001834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:24:53.001834Z digest=sha256:766f3c0595b6ce6224333fcd30b252c54ad9a32fd71b283e5a3c516226039ee8

Observation 01dc94d4-e247-4426-9290-60f5a4144ba9 · inbound

VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback? cites this paper.

VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback? A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets

Reference 179

Resolution
unresolved
no resolver link, observed 2026-08-15T14:25:40.316937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:25:40.316937Z digest=sha256:d71a8ef267d409792dc005940d43112d6584ab81d324831dc2d67040bc9b73f4