Pith. sign in

Paper Citation Record · LEDGER

How to Choose a Threshold for an Evaluation Metric for Large Language Models

As of 22 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2412.12148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12148 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:28:07.694891Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f3bb506b-d45a-49b6-bfac-48d4cf5cdddc · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:08.173739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.562968Z digest=sha256:5f8002b44f07fe4bbd540dfa24b0d425e157011c690a891b8445ba596e96267f

Observation 5e58eb0f-7d6b-43fd-8cf8-d42d5175667b · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:08.156813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.569952Z digest=sha256:8817bb4e68eb6dad490374e051ddf7673352f0e4d30711e48a270b50a87c1337

Observation 0ba46453-5a63-4fe7-aab1-0d792bef343e · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:08.138699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.576475Z digest=sha256:2a0ec6af308796d7fb3863797e40c962bb4e6fcc2a57d10a805a234cd07c295a

Observation b81dd38c-b15a-45c4-a8cf-67a4d587af63 · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:08.119292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.582534Z digest=sha256:348ac12890d25373aa463aaa6af42498c610dd9239cd447c45e304ec6d21b311

Observation 3fff6ead-01b6-4366-8863-cb38d428b2c5 · outbound

This paper cites Ragas: Automated Evaluation of Retrieval Augmented Generation.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Ragas: Automated Evaluation of Retrieval Augmented Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T18:28:07.588952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:28:07.588952Z digest=sha256:22efbfd5aeaaff6edee75f0d1b2cd5389f718527b044b8b312a1bdd65fcfec3f

Observation eb47dc57-1fd4-4f9c-b3c6-e33711b03902 · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T18:28:07.594727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:28:07.594727Z digest=sha256:45f5e5d56963d9a5cacf38bdbd2ff1dc2bc2d459f31b1f274cf418a7c676d358

Observation e02c04f2-5e3f-47ba-ac7c-89a83377571d · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:08.101417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.601099Z digest=sha256:5aaf5ca2bfc69c7f64c9ee7606e72eb432fa7e7c1318e3d4ba7bdcfcfff23ada

Observation 7296a313-9e75-436e-a812-d1d5c460db69 · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:08.083104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.606323Z digest=sha256:3e18b9b0900cc59b88e4f4ec316749d99fae8a22b5e9c268ece5ae2918a129dd

Observation d1f8baa9-04c5-432a-a968-d13a19e7ae84 · outbound

This paper cites Kahneman and A.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Kahneman and A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:28:08.065298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.611655Z digest=sha256:af2019fa69552fa5caf7556a63ece710e495fa1fb4c3478a72d2413ed25148b8

Observation 733bb6f1-130f-4143-ae1a-aebdeaee6fc4 · outbound

This paper cites Lewis, E.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Lewis, E

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:28:08.047617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.617773Z digest=sha256:5cd3228fabae3454fa13fa60e5dadf7e4e63c321c8a09caef0d6ca8adcf4a17f

Observation 23859db2-96ce-4c2a-9c71-cacfb39daed9 · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T18:28:07.623216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:28:07.623216Z digest=sha256:c65e33b0cc0e5ec5228cc0fb461dd0e1320fa65c043305eda3738f2640df9ad9

Observation 3b042821-393c-4f21-a6f7-f874576ab6c1 · outbound

This paper cites Schroedinger's Threshold: When the AUC doesn't predict Accuracy.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Schroedinger's Threshold: When the AUC doesn't predict Accuracy

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:28:07.830094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.628285Z digest=sha256:4e9f5e910ddf521cb69633fe43fb5d4864e9fdf043e2d95afb28b91a5feb509e

Observation 51b4cb5d-3ae5-497c-9dcd-6db824dc4ec9 · outbound

This paper cites Papineni, S.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Papineni, S

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:28:08.018154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.633789Z digest=sha256:19068723a6bbb4ea8e6f9cee97aafbeda36733835baa5d35f9b79a14b252320e

Observation 29857475-e69f-472f-8fdc-c0d4075d152f · outbound

This paper cites Lynx: An Open Source Hallucination Evaluation Model.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Lynx: An Open Source Hallucination Evaluation Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T18:28:07.639164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:28:07.639164Z digest=sha256:2ce749536494074fe5855e94af2d56356440bb0b495f3eed09e34e5b1ed2df49

Observation c496e97b-1cc0-4598-b4b7-8b769f8a70c7 · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:08.000120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.645363Z digest=sha256:6192cef2f38ee2daeb5501771cd3f4aa498181d2efdedcc28264200626eb1078

Observation 30ab993c-7640-4167-bdbc-bf61a8c24055 · outbound

This paper cites Sadinle, J.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Sadinle, J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:28:07.980494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.650669Z digest=sha256:a54d07354c5d4b2e3649c0cdc6b7da45d48f883c48596e62f2771f660f87ef40

Observation 53d526f2-0623-4d2f-8911-15cc65f2634e · outbound

This paper cites Sudjianto and S.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Sudjianto and S

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:28:07.961691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.655865Z digest=sha256:a1a2ff61ffa592f8c2caa86d983b76d07f4301cd7129453f5e982d3bf8d17084

Observation 6eb14151-9d32-47ab-9cbc-73ea88392d49 · outbound

This paper cites Model Validation Practice in Banking: A Structured Approach for Predictive Models.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Model Validation Practice in Banking: A Structured Approach for Predictive Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:28:07.781577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.660930Z digest=sha256:d11d9d1c1d31eb3ca784c21d731aca2685ac34cbb48a4c95c64e5d44122be130

Observation aba69b2f-b942-4418-bda1-c35b6067bebe · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:07.943796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.667204Z digest=sha256:0a29f5f868f654c633630e533ec64b95c7f8de77595d32cc0a2fb764f99585da

Observation b9689252-f223-464d-a931-69bec90dea76 · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:07.927197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.672619Z digest=sha256:c0211981e472497c6265f2997924606fac6907ff75797643f31d72f02570da6a

Observation 849bc6f1-dea6-46b8-8762-858f124b0665 · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:07.908865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.677809Z digest=sha256:6b2d013c239a18dff9f7a6653ebbafee668a29faf5e4b4cd8cc60420acdb8a58

Observation ffe5f2af-4d59-4f50-ae7e-1d1996852f6b · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

How to Choose a Threshold for an Evaluation Metric for Large Language Models BERTScore: Evaluating Text Generation with BERT

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T18:28:07.683123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:28:07.683123Z digest=sha256:04fada05a553fc3130b6abf3d13a262e686224f208909838e0bc6f65766c93b3

Observation f185a0e4-47bc-40d4-977c-439f6e5deff4 · outbound

This paper cites A Survey of Large Language Models.

How to Choose a Threshold for an Evaluation Metric for Large Language Models A Survey of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T18:28:07.688990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:28:07.688990Z digest=sha256:b22bc7fd135c0be85530a5a06b473b871378d8ff0d565d32085fe86d179b47a8

Observation f42eecd1-3171-41ef-8a36-eff732989fcc · outbound

This paper cites Figure 6 further illustrates the identified thresholds across different risk tolerances, measured by the false positive rate (type I error) and precision.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Figure 6 further illustrates the identified thresholds across different risk tolerances, measured by the false positive rate (type I error) and precision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:28:07.891001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T18:28:07.694891Z digest=sha256:b1707d98573f2cbe6a4a8a99915bf5168bff6a025d92b135849260ad15206f98

Pith citing papers

No inbound Pith citation observations are available.