Pith. sign in

Paper Citation Record · LEDGER

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models

As of 13 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2411.15320.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15320 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:29:54.284637Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T18:27:23.076544Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T18:31:44.524453Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6a8cb7d-5d4d-453d-af05-4bb025f95416 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.171039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.171039Z digest=sha256:d60f9daf6dad031523468529a01291f68e05a52d2e53b09fd3343a39e9db6d52

Observation af84ff0f-652b-464d-af39-d514e287583a · outbound

This paper cites write newline.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.176096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.176096Z digest=sha256:1ff3787e8db7ba4ca0b8f0962141bf2daac136671abdd897fdd50acde4740139

Observation 4f807f6a-add5-4a53-a64a-1e10b2e9877c · outbound

This paper cites Benchmarking Foundation Models with Language-Model-as-an-Examiner.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Benchmarking Foundation Models with Language-Model-as-an-Examiner

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.180727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.180727Z digest=sha256:044c7afd6ee5e6875db8e45bfb47ab268b46f650236ab195f328ef2a2c5e01f7

Observation f97974cb-c8d3-44c4-a663-cad246c47a33 · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.604747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.185421Z digest=sha256:2284b4c5087611fb91b63484b2bb81a4f53fe3ff7fc191eebbd93b5a4d6b3ee8

Observation 92ede92c-4fd9-40fe-afec-76154bfe11be · outbound

This paper cites J.; and Jurman, G.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models J.; and Jurman, G

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:29:54.591881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.189712Z digest=sha256:a7dd5a991146c140b1c1dc9890d4092d023b85eb953a25dafdea26ffdad9a74a

Observation 95851375-d731-4b6a-9afb-76c012183f61 · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.580303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.193802Z digest=sha256:f83ce14fe3641ca166e5391df6c3011c6904ba7402193250e365be466bd61e05

Observation 7f7695b5-5bea-4e40-a80e-11962a72be5f · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.569592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.197870Z digest=sha256:9bfcbca964039adf0902e363bf1d0d492b4438e061046cd52d9e094255f95d62

Observation 9862311d-7274-436e-a049-4d33d266bacb · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.557458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.201945Z digest=sha256:8a01a078f63a0508566917f2f2e7d7ac8d7bc0b67266aa938f92bc9f521916a6

Observation 5736b260-f1f9-4fb1-a2ab-193f4afdaf42 · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.545583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.206207Z digest=sha256:0419dd82b5af411b09a866e691c2c435d2e6d34b75c0c92cad5dcbcdbda8e956

Observation eca368e3-d7ab-4f54-ace5-abb1cacc6d0b · outbound

This paper cites GPTScore: Evaluate as You Desire.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models GPTScore: Evaluate as You Desire

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.210042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.210042Z digest=sha256:7d6fb0929a2cbd57ca89e42995c4c1ba65a5a57aaa106d0d93d78d6653c530bd

Observation 51f47c82-cf10-4bf2-98dd-215352389f4d · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.214139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.214139Z digest=sha256:666fd29516bc7c4c5164eb210fd948c4db2f911bca0700a11f6b496d41bcfd67

Observation 6629d659-0bc5-404a-83cd-4b4a290c9877 · outbound

This paper cites Demystifying Prompts in Language Models via Perplexity Estimation.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Demystifying Prompts in Language Models via Perplexity Estimation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.218118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.218118Z digest=sha256:26aa015c3783c30cf6444df434f3317d1fe322efebe13d24b00f42ce02da0d18

Observation c6bcbf22-ed82-44d4-a177-9ab5992b46aa · outbound

This paper cites OLMES: A Standard for Language Model Evaluations.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models OLMES: A Standard for Language Model Evaluations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.222181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.222181Z digest=sha256:dc8c521c0cb91c060bd79fa9b211abade027e86c2048b6e85df88de98d69c397

Observation 45e17965-2ffe-431c-bd3b-2ba97c179f1d · outbound

This paper cites Holistic Evaluation of Language Models.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Holistic Evaluation of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.226325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.226325Z digest=sha256:7617f352efe046beef051af765c31d2e6c5ef61dd9df5f0b1fe18cfc8f00112b

Observation d0ec3e1e-495e-4972-8c14-7cd3f979194f · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.526086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.230138Z digest=sha256:72b189cc9b996b0321fed5c96fe095f59766ec76d9288c0146571a96f47494c7

Observation 0b709059-0e66-46cb-adbc-02615707342e · outbound

This paper cites LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.233656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.233656Z digest=sha256:926071eb5908a6013f346c2326f88130d14165111944d0eb72b3b2811a1a15ff

Observation e2fbde52-920a-424f-b224-4d129ac0a1a6 · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.514068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.237459Z digest=sha256:0f1192c4eae737a7c21a830addba41f92ff28633d941f85c0c6181559facd593

Observation c37fd807-f80c-4f3c-a7b2-a4f2510030dd · outbound

This paper cites GPT-4 Technical Report.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.241101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.241101Z digest=sha256:8237bf753e4a6e17ab7d9a6944a445c31af79b2904d8a59197d780ee02f1e5c6

Observation b80bfecf-3dca-4998-bdb8-1e6f73568866 · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.502295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.245111Z digest=sha256:5128ebddcd079d847140d7f54032d4a7eb2275e0cea2dffa8983bfb0b7fdb2c2

Observation 68deed28-a79c-4e4a-9e8e-b625b5523a9a · outbound

This paper cites Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.248917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.248917Z digest=sha256:7feb5ad4f03137585a0c30da5c00564f39b997393cd0907d89b045ffab7f18b6

Observation 1c4e7523-582a-4ae9-a987-67840539f7b4 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.253765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.253765Z digest=sha256:bf8f12ba6c1b3db16b3d58fdde25f9c3f7466717c75b664f9872d41e3e5372b4

Observation afc954c1-e2ff-4f77-9be5-60d365c448c6 · outbound

This paper cites L.; and Parikh, D.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models L.; and Parikh, D

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:29:54.491310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.257663Z digest=sha256:9ded64e5bb015e7fd8cc6dcf29ed6cce67f3488d0259e5e613e235419bd1cc4b

Observation ec311917-7048-4f14-a72a-bea9dfee9684 · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.479863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.261221Z digest=sha256:16fbedfa4e058468a7a95cd04f67640be5353a1f21088985ec1971ce1c902c52

Observation fd272bb9-6ebb-4a7b-aa4e-7e87dcc15d8e · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.467074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.264645Z digest=sha256:249a31c5c8100bb2ba3488d3c820c60d3b5d7529f28479f8ff08bbc260042219

Observation 068c3815-3ce9-4658-9630-25aa6eacc72c · outbound

This paper cites an unresolved cited work.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:29:54.454739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T14:29:54.268688Z digest=sha256:477556d83c0be30537f8dad15030718de12b9f760036edefe975b065a6a7713d

Observation 3aa9fed1-7440-45ca-b0bd-3b2d6da62464 · outbound

This paper cites A Systematic Evaluation of Large Language Models of Code.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models A Systematic Evaluation of Large Language Models of Code

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.272514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.272514Z digest=sha256:dd625a2351f172c865db2c459a62e7bd0627221d2181456c719cfa4be9f8406d

Observation 05efc8a0-510d-4f6a-b0ff-b90056466895 · outbound

This paper cites EvalAI: Towards Better Evaluation Systems for AI Agents.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models EvalAI: Towards Better Evaluation Systems for AI Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.276421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.276421Z digest=sha256:65bbdaec14dc8ca5fca5558a95302a1d0542b89c2c69bc3976c0a5eaa9e74868

Observation 333d1227-68fb-4d2e-8c22-969d030b4255 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models BERTScore: Evaluating Text Generation with BERT

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.280278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.280278Z digest=sha256:c1876899debbd42661c68b4dd9b7693202b369ce8720f8a3da85de4e03e02395

Observation f21dfe25-205d-4a58-be6d-cef4100443c4 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:29:54.284637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:29:54.284637Z digest=sha256:15398346213024990e2d92a496bed52242ae27cb0180bdd30ecd6d71cfb59bf7

Pith citing papers

Observation 59e10ad5-fa82-4eba-8981-577f18095a8d · inbound

Self-Aligned Reward: Towards Effective and Efficient Reasoners cites this paper.

Self-Aligned Reward: Towards Effective and Efficient Reasoners PPLqa: An Unsupervised Information-Theoretic Quality Metric for Comparing Generative Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.527457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-18T18:27:23.076544Z digest=sha256:812c0e5d15313d93aaebb3319bc51dd98305ecedddedfd41071545d104a6ad49