Pith. sign in

Paper Citation Record · LEDGER

Is ChatGPT a Good NLG Evaluator? A Preliminary Study

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2303.04048.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.04048 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:44:21.334991Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0cd08ba4-c0c0-4b19-97d5-f78d227a9fff · inbound

G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment cites this paper.

G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T22:55:50.901173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T22:55:50.014540Z digest=sha256:9fccf48d617e7d6846e8fcee0182a256fe2cf9dd61f160f7fac27a7a7521fee4

Observation 2a945c7e-7d2a-48b2-bcd2-39f3bffba0e4 · inbound

ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate cites this paper.

ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:03:18.849944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T13:03:18.765496Z digest=sha256:323b0d987f072c279c4710065d48f15dd4ec3763caceafbaf76f2d081a993b40

Observation da3b69b4-b1ad-43a3-a649-ca760b30037e · inbound

GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts cites this paper.

GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:25:21.097589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T06:25:20.966510Z digest=sha256:2d074d03e1faf4f2d941b6354963a700af1dfa0678593181246a10c70d7893f1

Observation d53759e9-017d-4afd-a1b8-aa63f2e7c5af · inbound

From Local to Global: A Graph RAG Approach to Query-Focused Summarization cites this paper.

From Local to Global: A Graph RAG Approach to Query-Focused Summarization Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:58.035351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T05:10:57.816312Z digest=sha256:a611ae4cc369b691581ae2c0ef3d3a3e9948349d2fc8531e3458c4f33f3fd8c6

Observation bc4e9816-f1ee-4c3e-8b68-81dca979c415 · inbound

Dictionary Insertion Prompting for Multilingual Reasoning on Multilingual Large Language Models cites this paper.

Dictionary Insertion Prompting for Multilingual Reasoning on Multilingual Large Language Models Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:05:44.746639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-23T18:03:19.212619Z digest=sha256:148efbe686e0dd8a65330d35ace7e67a11563b9ab1cdbc4701f006bd068aba80

Observation fb4babe8-2286-4312-9de9-f81bf65a5075 · inbound

In-depth Analysis of Graph-based RAG in a Unified Framework cites this paper.

In-depth Analysis of Graph-based RAG in a Unified Framework Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:37:22.326507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:36:25.057478Z digest=sha256:ff0277a0649b4ff5eb520fea185d035c7743d246a5edae5b09ae5db0fecd0a86

Observation 7074985d-727f-467a-9264-d029930fb6de · inbound

SLoW: Select Low-frequency Words! Automatic Dictionary Selection for Translation on Large Language Models cites this paper.

SLoW: Select Low-frequency Words! Automatic Dictionary Selection for Translation on Large Language Models Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:34:26.392442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T23:31:58.225436Z digest=sha256:6d6a486b29010a242957fd265b2527fc7d93ab18a7df8106f2f34207332debc3

Observation 6b9d5de5-5b61-4411-aa16-363f63b34525 · inbound

Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts cites this paper.

Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T15:44:21.334991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:44:21.334991Z digest=sha256:ea09ea2c200aa5de117503c3623f2aa78befcaa464b5bc95db38b888ea29191a

Observation 0a3412f4-d392-46eb-9f58-4b818546cf79 · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:31.375425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:31.375425Z digest=sha256:7cf81cf2ee45b4f01dca516910d8faf1b63eabdb67d31d27a5f48c51993b435b

Observation 37094538-b6cc-4d56-a730-73334d089f0a · inbound

RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model cites this paper.

RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T11:32:03.720711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:32:03.720711Z digest=sha256:fb913767753aa65de35001f343b183760f99f8e0b1ec7702c081856fcc7459fe

Observation 0c864c14-98df-4288-ae5e-7a1bfb264827 · inbound

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data cites this paper.

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:49.943824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:49.943824Z digest=sha256:d282a236158ae44caac987d05ff945322c7d700c3bd398be4498a2ae9090353c

Observation d11aba2e-8ed6-4286-91e1-cc67ff93c667 · inbound

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method cites this paper.

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:42.132070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:42.132070Z digest=sha256:63db66ab80d5d13df6b9f153204f16abf86999826d763fda75719a9998c0da71

Observation f1f30f70-58da-4bda-9248-7d40dbb60e40 · inbound

PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses cites this paper.

PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:00:03.058822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T13:57:41.428695Z digest=sha256:da45094d60b39edc1e70a83eb536614913a1c929b49fa68e08877c41cb039b16

Observation 5d918467-b1be-45b2-9bde-d2798909e7d0 · inbound

MMP-Refer: Multimodal Path Retrieval-augmented LLMs For Explainable Recommendation cites this paper.

MMP-Refer: Multimodal Path Retrieval-augmented LLMs For Explainable Recommendation Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:33:02.616446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T17:28:28.480792Z digest=sha256:922eac72d6c6461e741cad3c42774f779465560c4ececb1c4b12f1452a9b8f99

Observation 30b93ae1-88e9-4b1c-9097-577eed10007a · inbound

Supporting System Testing with a Multi-Agent LLM-based Framework for Knowledge Graph Extraction: A Case Study with Ethernet Switch Systems cites this paper.

Supporting System Testing with a Multi-Agent LLM-based Framework for Knowledge Graph Extraction: A Case Study with Ethernet Switch Systems Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:33:08.873287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T08:32:39.748993Z digest=sha256:bb72255547165df51fbd649425f03b2a79ae91add122ef8848bbd3a9d5cffadb

Observation cf66b034-f40b-41cb-a939-c61f338d46b9 · inbound

A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models cites this paper.

A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:22:37.605994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T19:53:36.990878Z digest=sha256:8402c903b1633ec55fe0a710ebbe8d71ab4999038674ace4e2689e4481201538