Pith. sign in

Paper Citation Record · LEDGER

Human-Calibrated Automated Testing and Validation of Generative Language Models

As of 13 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2411.16391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16391 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:13:17.824953Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T06:34:54.572363Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T06:36:43.165501Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d475b582-1108-4126-afba-4325da029882 · outbound

This paper cites an unresolved cited work.

Human-Calibrated Automated Testing and Validation of Generative Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:13:18.514309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.672139Z digest=sha256:c6f4214bf196b693933f250fec299ac744efadfb63a8c79d436babff7360b68b

Observation f4e9db51-8048-48bd-81f7-79874e3db957 · outbound

This paper cites D., Dhariwal, P.,.

Human-Calibrated Automated Testing and Validation of Generative Language Models D., Dhariwal, P.,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.486573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.678908Z digest=sha256:b0f7060567bf232ead28f977b349ffbd5e38e6fee75d3cf47198107f1b9dea3e

Observation 18cd92b3-c97c-4d80-98cd-ff469246b953 · outbound

This paper cites Statistical optimal transport.

Human-Calibrated Automated Testing and Validation of Generative Language Models Statistical optimal transport

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.685051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.685051Z digest=sha256:1252f19d6c40390ed01002b303271c503db79750e491ee5249a47e07cf0f5128

Observation 73125079-3cb5-4ff4-bf98-5705f4736c16 · outbound

This paper cites W., Lee, K.

Human-Calibrated Automated Testing and Validation of Generative Language Models W., Lee, K

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.463189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.691030Z digest=sha256:0f9a1216beac453c59a01383f616c934aaf5ce1780c8ec642e799193ffe67d52

Observation 9829d707-54cf-4392-be4b-13c53b976abb · outbound

This paper cites and Chen, D.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Chen, D

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.442102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.697137Z digest=sha256:4fb7b74d54a5a99209749df34c430a03dc3aefcfe602fe16618561bd9817ac6b

Observation 79110c0c-b98b-4f87-8de6-d018d239278d · outbound

This paper cites BERTopic: Neural topic modeling with a class-based TF-IDF procedure.

Human-Calibrated Automated Testing and Validation of Generative Language Models BERTopic: Neural topic modeling with a class-based TF-IDF procedure

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.705744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.705744Z digest=sha256:69f9cdb0af0b08c62ea23240554975c1426b8a1f7e0c7cb61d9252580e01a700

Observation f29a58c7-0128-4e4c-a7fc-9c186e572ecc · outbound

This paper cites and Unitary team.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Unitary team

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.421197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.712305Z digest=sha256:e2d6b19e7516e088dbabb9e5f0bff79cac205a1b031e99095434f4cf69fc2919

Observation f16d3497-1404-4d1c-abd1-b4c3790b73ac · outbound

This paper cites QAFactEval: Improved QA-Based Factual Consistency Evaluation for Summarization.

Human-Calibrated Automated Testing and Validation of Generative Language Models QAFactEval: Improved QA-Based Factual Consistency Evaluation for Summarization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.717427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.717427Z digest=sha256:6aa0442413c736130105cbd5983f94b26d1fcd236b861350ba9c32f99e275e96

Observation 4fd2f648-8b76-4966-a9f2-d8d820dc4b68 · outbound

This paper cites Asking and Answering Questions to Evaluate the Factual Consistency of Summaries.

Human-Calibrated Automated Testing and Validation of Generative Language Models Asking and Answering Questions to Evaluate the Factual Consistency of Summaries

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.723487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.723487Z digest=sha256:4ed1cfa5740143b3937c4eed6a76494cf2fa418fcd42b15bff7c6e1f3f246b2b

Observation 7f4a565d-eaee-4efa-872f-9024cd6dd397 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Human-Calibrated Automated Testing and Validation of Generative Language Models Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.729889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.729889Z digest=sha256:f87acca9a4c8cec3f5286cc21effc6ee5c66076253aec1c4ded636591b349e38

Observation bb7369d6-7286-4e19-b069-69e2639c038b · outbound

This paper cites and Bansal, M.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Bansal, M

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.394189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.736659Z digest=sha256:2edc651849d6b2c42794643972a93e688a038809fa7b563d7b1f1e919aff8a3b

Observation 374488d9-b290-4812-a200-40547505b47b · outbound

This paper cites , and Kiela, D.

Human-Calibrated Automated Testing and Validation of Generative Language Models , and Kiela, D

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.370811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.741914Z digest=sha256:afa34545af42e8dbf5fc67a598293f1a21885c39724d5918ea28957c443fbb36

Observation 3e2c512d-5986-47aa-a197-d6837d1f50b2 · outbound

This paper cites , and Koreeda, Y.

Human-Calibrated Automated Testing and Validation of Generative Language Models , and Koreeda, Y

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.349617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.746694Z digest=sha256:5bd30a99ea1ae7ca304670daf1477e4053f8fc1f0122f777aa996db89f30751b

Observation fc479bc5-748b-4df0-b67b-67e5d5fc5843 · outbound

This paper cites and Evans, O.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Evans, O

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.323646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.752246Z digest=sha256:b8e9da8d13704b87ee14ac7538876314e857ea8f59f4063414ac6d2529926fc5

Observation 67de1eb9-3f2e-415d-a994-db6e53278f2b · outbound

This paper cites Automatic Generation of Behavioral Test Cases For Natural Language Processing Using Clustering and Prompting.

Human-Calibrated Automated Testing and Validation of Generative Language Models Automatic Generation of Behavioral Test Cases For Natural Language Processing Using Clustering and Prompting

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:13:17.929626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.757606Z digest=sha256:9fe214967c5dd090342bc022e52e43b5367c6ef37b021aaa848a845b1bcd8553

Observation caee6645-1db0-46cd-a39c-1fc25ec914c1 · outbound

This paper cites an unresolved cited work.

Human-Calibrated Automated Testing and Validation of Generative Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:13:18.284067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.762900Z digest=sha256:074cac6e9164c666ebd2b876c8b90638aaaaa2576424e853d57ad53587ca8a00

Observation b94ccbd7-7136-4986-8e96-7986c982258c · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

Human-Calibrated Automated Testing and Validation of Generative Language Models UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.768083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.768083Z digest=sha256:d30e8c9fa648799264446a9fcf78553c9e06e7b9efae6519d48826388645012c

Observation d8033f59-25ba-4b49-9212-4e41eee2da74 · outbound

This paper cites and Gurevych, I.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Gurevych, I

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.261978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.773880Z digest=sha256:1247628f386e29f552bd9fa6184212ab596ebf50789f274913a90555b8ba4bc5

Observation 5d3a9995-e779-420f-82be-5e838272b0e4 · outbound

This paper cites T., Wu, T., Guestrin, C.

Human-Calibrated Automated Testing and Validation of Generative Language Models T., Wu, T., Guestrin, C

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.238128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.778866Z digest=sha256:2b9a8256035ea094a225e03464c9af2a6f43e57878fc1aa0e998f603b1e80bb8

Observation c3a03187-3093-4985-a3f8-719fca11cecd · outbound

This paper cites P., and Xu, X.

Human-Calibrated Automated Testing and Validation of Generative Language Models P., and Xu, X

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.213127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.785378Z digest=sha256:c9fdeeeb70fbc0c0a970c7d683d7616285d4360afdb017661cf33902b67e1ee3

Observation 2da941b6-e1ba-4601-8468-6df268d54591 · outbound

This paper cites and Zhang, A.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Zhang, A

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.187800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.791062Z digest=sha256:d42be2a8639ab01365a58655032098873494fa824ebbc7ac0f8174e50d97ddd7

Observation f7ee1654-0631-4eff-a88a-a482e2243345 · outbound

This paper cites A., Abid, A., Fisch, A.,.

Human-Calibrated Automated Testing and Validation of Generative Language Models A., Abid, A., Fisch, A.,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.156961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.797762Z digest=sha256:1658fa3f7ce8e3dc42b366a47401f3ee7b15036bbe169c0159fe8bcf729dd43f

Observation 475f1107-1229-4c5c-8d0d-a64d4b21caf8 · outbound

This paper cites and Wang, Z.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Wang, Z

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.129505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.803096Z digest=sha256:6b23739b95c229af9a5ddea959e7326edccd8678bdb5fe87bbadab34070bb04d

Observation ffb5d2f9-2f01-4a0b-ac99-5ad382ad7b3c · outbound

This paper cites and Shafer, G.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Shafer, G

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.103309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.808402Z digest=sha256:032a90362097345ad9d547f0c2386dc189388710c1519996732bd111c5725dcf

Observation 834ad3c5-9147-4512-bac3-2a5cbc712681 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Human-Calibrated Automated Testing and Validation of Generative Language Models HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.813782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.813782Z digest=sha256:67f2f984b9593b10871f62c47bd4122c30ceddff14694eaed45d221989ec22b1

Observation 04f7eaa7-a087-4098-a8bc-fb02e429a0a8 · outbound

This paper cites an unresolved cited work.

Human-Calibrated Automated Testing and Validation of Generative Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:13:18.073918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.819874Z digest=sha256:507c80e67063d2335d06db2170ac33817bf878545fb4e205f3e38bd3d9ca4c53

Observation d0e46e76-711b-4f72-ad47-c0db79709e49 · outbound

This paper cites A Survey of Large Language Models.

Human-Calibrated Automated Testing and Validation of Generative Language Models A Survey of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.824953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.824953Z digest=sha256:6847e14fc61de23004b845bac34b741483754d7b5bace4cd4217dbd4fbec5b6b

Pith citing papers

Observation 19fe34e4-4d2f-4d07-ac69-b5e58bd3892a · inbound

As It Was: Aligning LLM Search Evaluation with Historical User Preferences cites this paper.

As It Was: Aligning LLM Search Evaluation with Historical User Preferences Human-Calibrated Automated Testing and Validation of Generative Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:36:43.167106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-02T06:34:54.572363Z digest=sha256:5a168cde329114bee3f468af9c5db288a9a225300f67a20978d4d47e47dae2ad