Pith. sign in

Paper Citation Record · LEDGER

How to Choose a Threshold for an Evaluation Metric for Large Language Models

As of 18 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2412.12148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12148 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:28:07.694891Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f3bb506b-d45a-49b6-bfac-48d4cf5cdddc · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:08.173739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.562968Z digest=sha256:18a9217edfa124b9a1508fd61665ac5e05af319dc6bc745935e49c0a5fa43c9d

Observation 5e58eb0f-7d6b-43fd-8cf8-d42d5175667b · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:08.156813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.569952Z digest=sha256:9079bb595dbb6667b2856c9bbbbcd9147fee60af4fda28060c1269266b7b4b82

Observation 0ba46453-5a63-4fe7-aab1-0d792bef343e · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:08.138699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.576475Z digest=sha256:806845752b1872d653c2bb278cfabb6f5271106a842c7bebd64430f5910f2fdd

Observation b81dd38c-b15a-45c4-a8cf-67a4d587af63 · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:08.119292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.582534Z digest=sha256:785131a2c9bfc8492df9de86b5bdd0148c7a24b2a44ff640d8f3c5452bedd696

Observation 3fff6ead-01b6-4366-8863-cb38d428b2c5 · outbound

This paper cites Ragas: Automated Evaluation of Retrieval Augmented Generation.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Ragas: Automated Evaluation of Retrieval Augmented Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T18:28:07.588952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:28:07.588952Z digest=sha256:da051fd3cb3f9c14ba1272294493971c5d6b222bc50cde6679adb2a1679f5c82

Observation eb47dc57-1fd4-4f9c-b3c6-e33711b03902 · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T18:28:07.594727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:28:07.594727Z digest=sha256:e43f915b8fc47c653e895dcd77098b29b63da19a507e72daf97d866f6329e16f

Observation e02c04f2-5e3f-47ba-ac7c-89a83377571d · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:08.101417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.601099Z digest=sha256:b3dd2790b53ab8ceed2d01b059104360437212d0d76caa0a856bac97b37b6db0

Observation 7296a313-9e75-436e-a812-d1d5c460db69 · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:08.083104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.606323Z digest=sha256:eed24fd6f405b5ebb34799baae9463cf8d24c94adf2b244448301f0f595e0476

Observation d1f8baa9-04c5-432a-a968-d13a19e7ae84 · outbound

This paper cites Kahneman and A.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Kahneman and A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:28:08.065298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.611655Z digest=sha256:7782e6c65cb1b2207635d108a2df58fcf023cc1fa4d79a71c7273f8304272448

Observation 733bb6f1-130f-4143-ae1a-aebdeaee6fc4 · outbound

This paper cites Lewis, E.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Lewis, E

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:28:08.047617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.617773Z digest=sha256:588936c42d736655e88e8ef0cffc3b911e8967edd83ca9dac1aba5b79185d49c

Observation 23859db2-96ce-4c2a-9c71-cacfb39daed9 · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T18:28:07.623216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:28:07.623216Z digest=sha256:d10524b0ade38f045092c780189de28c13433095fb60b5b2eb2be7c85a7c4b74

Observation 3b042821-393c-4f21-a6f7-f874576ab6c1 · outbound

This paper cites Schroedinger's Threshold: When the AUC doesn't predict Accuracy.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Schroedinger's Threshold: When the AUC doesn't predict Accuracy

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:28:07.830094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.628285Z digest=sha256:4d872ce67352d8a1ea43d850de38893f00d6e3e799e8ee4a323b1456db3ecbba

Observation 51b4cb5d-3ae5-497c-9dcd-6db824dc4ec9 · outbound

This paper cites Papineni, S.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Papineni, S

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:28:08.018154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.633789Z digest=sha256:04675d7cad4b3585115a5e82fa700ff0a61a6cbdc9830ccc2f87b5f8726003c8

Observation 29857475-e69f-472f-8fdc-c0d4075d152f · outbound

This paper cites Lynx: An Open Source Hallucination Evaluation Model.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Lynx: An Open Source Hallucination Evaluation Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T18:28:07.639164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:28:07.639164Z digest=sha256:e325e9e2bb8b30d8870ce1c4606557590e0cbc71c5337bc14f22c1c9907c31f8

Observation c496e97b-1cc0-4598-b4b7-8b769f8a70c7 · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:08.000120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.645363Z digest=sha256:a023de0fc53318afbdb8cbdbc7041dd279fabd061402c31d40cffd2512840e67

Observation 30ab993c-7640-4167-bdbc-bf61a8c24055 · outbound

This paper cites Sadinle, J.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Sadinle, J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:28:07.980494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.650669Z digest=sha256:4dc69c5b0ad87974b790237a2dac69e994e060cb0dd8b53a06e64e4a3d31d517

Observation 53d526f2-0623-4d2f-8911-15cc65f2634e · outbound

This paper cites Sudjianto and S.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Sudjianto and S

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:28:07.961691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.655865Z digest=sha256:e35901b7fde6e1e21cf4137a167334acfdc0d896c5f82446770bdc7ca4974884

Observation 6eb14151-9d32-47ab-9cbc-73ea88392d49 · outbound

This paper cites Model Validation Practice in Banking: A Structured Approach for Predictive Models.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Model Validation Practice in Banking: A Structured Approach for Predictive Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:28:07.781577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.660930Z digest=sha256:3c8e9cf08db54e69585e26ac27224764aca85410021d0497cf310e5014806e60

Observation aba69b2f-b942-4418-bda1-c35b6067bebe · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:07.943796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.667204Z digest=sha256:7fc864fb8e25015932a0b0eef287244129e4fbe4a3d9fb94094463c29b10ee5f

Observation b9689252-f223-464d-a931-69bec90dea76 · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:07.927197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.672619Z digest=sha256:965d17e5cb6d6a1e92fccc7503f46d960cee13270f5b71360db4d35f80fa2cdf

Observation 849bc6f1-dea6-46b8-8762-858f124b0665 · outbound

This paper cites an unresolved cited work.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:28:07.908865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.677809Z digest=sha256:acb708994b722925b8823ecde8b718253ff073e939e7f4d7dc373f57635311f0

Observation ffe5f2af-4d59-4f50-ae7e-1d1996852f6b · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

How to Choose a Threshold for an Evaluation Metric for Large Language Models BERTScore: Evaluating Text Generation with BERT

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T18:28:07.683123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:28:07.683123Z digest=sha256:b33f8cccc6bc171fc0c6ac7333d951648fb624f976425e0ab2ee4d9f7ab466bf

Observation f185a0e4-47bc-40d4-977c-439f6e5deff4 · outbound

This paper cites A Survey of Large Language Models.

How to Choose a Threshold for an Evaluation Metric for Large Language Models A Survey of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T18:28:07.688990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:28:07.688990Z digest=sha256:114d726100b2b1ba7709450fdb08f10a30faee66bd7ee9634859bd7bcbc9ac05

Observation f42eecd1-3171-41ef-8a36-eff732989fcc · outbound

This paper cites Figure 6 further illustrates the identified thresholds across different risk tolerances, measured by the false positive rate (type I error) and precision.

How to Choose a Threshold for an Evaluation Metric for Large Language Models Figure 6 further illustrates the identified thresholds across different risk tolerances, measured by the false positive rate (type I error) and precision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:28:07.891001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T18:28:07.694891Z digest=sha256:d1c45ae89291a8f9ca51cbed218fce3000aa35f77a532448d51681cace4b8c75

Pith citing papers

No inbound Pith citation observations are available.