Pith. sign in

Paper Citation Record · LEDGER

Human-Calibrated Automated Testing and Validation of Generative Language Models

As of 13 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2411.16391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16391 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:13:17.824953Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T06:34:54.572363Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T06:36:43.165501Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d475b582-1108-4126-afba-4325da029882 · outbound

This paper cites an unresolved cited work.

Human-Calibrated Automated Testing and Validation of Generative Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:13:18.514309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.672139Z digest=sha256:0ca764b435d70bed4e928b99996439c1e1af132855b908b0ba520c5a3e37febb

Observation f4e9db51-8048-48bd-81f7-79874e3db957 · outbound

This paper cites D., Dhariwal, P.,.

Human-Calibrated Automated Testing and Validation of Generative Language Models D., Dhariwal, P.,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.486573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.678908Z digest=sha256:0c7e6bae3f945b790fd2156224dc3fdbdc747214e4ab74e19233752fd6257766

Observation 18cd92b3-c97c-4d80-98cd-ff469246b953 · outbound

This paper cites Statistical optimal transport.

Human-Calibrated Automated Testing and Validation of Generative Language Models Statistical optimal transport

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.685051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.685051Z digest=sha256:668835585168340e44d52b2b4c7e35990abcf8b816df82b2ed7311373cad9498

Observation 73125079-3cb5-4ff4-bf98-5705f4736c16 · outbound

This paper cites W., Lee, K.

Human-Calibrated Automated Testing and Validation of Generative Language Models W., Lee, K

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.463189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.691030Z digest=sha256:de47995fcc11dd1eda30648800bcae6026a57f878bf09353162246823d2c649c

Observation 9829d707-54cf-4392-be4b-13c53b976abb · outbound

This paper cites and Chen, D.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Chen, D

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.442102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.697137Z digest=sha256:0fffa12e19310e357e8655beada93040f901081e08f3d364bb739559d71f63c2

Observation 79110c0c-b98b-4f87-8de6-d018d239278d · outbound

This paper cites BERTopic: Neural topic modeling with a class-based TF-IDF procedure.

Human-Calibrated Automated Testing and Validation of Generative Language Models BERTopic: Neural topic modeling with a class-based TF-IDF procedure

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.705744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.705744Z digest=sha256:36555984e38029c6f4e98ab3d90441c7b967f32085ee8bd19fa36be2033e8d2f

Observation f29a58c7-0128-4e4c-a7fc-9c186e572ecc · outbound

This paper cites and Unitary team.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Unitary team

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.421197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.712305Z digest=sha256:64788fdcf64ec12244475d809da2206ac268b4048720903d63f32c73981aee24

Observation f16d3497-1404-4d1c-abd1-b4c3790b73ac · outbound

This paper cites QAFactEval: Improved QA-Based Factual Consistency Evaluation for Summarization.

Human-Calibrated Automated Testing and Validation of Generative Language Models QAFactEval: Improved QA-Based Factual Consistency Evaluation for Summarization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.717427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.717427Z digest=sha256:d531cf3734a2e19a1415c1b3b6dbf48173fd1a52ec3a254175b1fc9e9be4d530

Observation 4fd2f648-8b76-4966-a9f2-d8d820dc4b68 · outbound

This paper cites Asking and Answering Questions to Evaluate the Factual Consistency of Summaries.

Human-Calibrated Automated Testing and Validation of Generative Language Models Asking and Answering Questions to Evaluate the Factual Consistency of Summaries

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.723487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.723487Z digest=sha256:0e5d742cb966711cbd340b17ac57431d5b5d8894892ba3c5b03d866478f1d6f1

Observation 7f4a565d-eaee-4efa-872f-9024cd6dd397 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Human-Calibrated Automated Testing and Validation of Generative Language Models Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.729889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.729889Z digest=sha256:1590ba9495380004116911b2c1e29c635a722cb47c0a78e96db18f630e4a51a9

Observation bb7369d6-7286-4e19-b069-69e2639c038b · outbound

This paper cites and Bansal, M.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Bansal, M

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.394189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.736659Z digest=sha256:a77691ee05bb3f3d154fb571de96af4398deb9240d2074b7530fa5ad278a8a60

Observation 374488d9-b290-4812-a200-40547505b47b · outbound

This paper cites , and Kiela, D.

Human-Calibrated Automated Testing and Validation of Generative Language Models , and Kiela, D

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.370811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.741914Z digest=sha256:daac809842b353219fee1a62189f0e1cc59a7ebe92a70c43d08627853a295bf1

Observation 3e2c512d-5986-47aa-a197-d6837d1f50b2 · outbound

This paper cites , and Koreeda, Y.

Human-Calibrated Automated Testing and Validation of Generative Language Models , and Koreeda, Y

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.349617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.746694Z digest=sha256:59971eed163f98c945ddec35aca67e9cbb4683a9f758a8edbbd5d7dc0d4b0562

Observation fc479bc5-748b-4df0-b67b-67e5d5fc5843 · outbound

This paper cites and Evans, O.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Evans, O

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.323646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.752246Z digest=sha256:ef43c0d172d6f5f3a6b6f16ed0451f189d5dee74450efd04d92091fbdb3e3fd1

Observation 67de1eb9-3f2e-415d-a994-db6e53278f2b · outbound

This paper cites Automatic Generation of Behavioral Test Cases For Natural Language Processing Using Clustering and Prompting.

Human-Calibrated Automated Testing and Validation of Generative Language Models Automatic Generation of Behavioral Test Cases For Natural Language Processing Using Clustering and Prompting

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-12T13:13:17.929626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.757606Z digest=sha256:7c5b75de474dfd090e26dc7f6dba9858bebece0176c0ac29f4cbd5d67de781f9

Observation caee6645-1db0-46cd-a39c-1fc25ec914c1 · outbound

This paper cites an unresolved cited work.

Human-Calibrated Automated Testing and Validation of Generative Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:13:18.284067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.762900Z digest=sha256:648a9160270e793d83cb697bb46da6484250e505bdafc9ae7ebded4ebe657951

Observation b94ccbd7-7136-4986-8e96-7986c982258c · outbound

This paper cites UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction.

Human-Calibrated Automated Testing and Validation of Generative Language Models UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.768083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.768083Z digest=sha256:01c0035a249a4ecc45cfc8f9f3787fc5799c3487bad7bac7e48864efcf86885d

Observation d8033f59-25ba-4b49-9212-4e41eee2da74 · outbound

This paper cites and Gurevych, I.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Gurevych, I

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.261978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.773880Z digest=sha256:00e8350e0ce1ed0f8dff48d26c3418fb52cbef370050189831fcd43dc633eafd

Observation 5d3a9995-e779-420f-82be-5e838272b0e4 · outbound

This paper cites T., Wu, T., Guestrin, C.

Human-Calibrated Automated Testing and Validation of Generative Language Models T., Wu, T., Guestrin, C

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.238128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.778866Z digest=sha256:d3c36e00197d45b2a1e6dbcbd288a42d23beb65953dbcd4b7dd82db20594b3a4

Observation c3a03187-3093-4985-a3f8-719fca11cecd · outbound

This paper cites P., and Xu, X.

Human-Calibrated Automated Testing and Validation of Generative Language Models P., and Xu, X

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.213127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.785378Z digest=sha256:3a7368324a8029cb7bb6bf75da9879b5d2d0fce4e5a1e0a9139c894a2e15e775

Observation 2da941b6-e1ba-4601-8468-6df268d54591 · outbound

This paper cites and Zhang, A.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Zhang, A

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.187800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.791062Z digest=sha256:5947931340aea29295f74829068848c9de9e3e246ef869e72d22c78343065398

Observation f7ee1654-0631-4eff-a88a-a482e2243345 · outbound

This paper cites A., Abid, A., Fisch, A.,.

Human-Calibrated Automated Testing and Validation of Generative Language Models A., Abid, A., Fisch, A.,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.156961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.797762Z digest=sha256:2f4f44ec469865c9d15291d172fb17a8c062d04b96a138b487874ecbaf93bdf5

Observation 475f1107-1229-4c5c-8d0d-a64d4b21caf8 · outbound

This paper cites and Wang, Z.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Wang, Z

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.129505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.803096Z digest=sha256:bbcac96d228da4e3783e9bc650aec76d40e1f04291e4c16a3e2a9cf8043d55d8

Observation ffb5d2f9-2f01-4a0b-ac99-5ad382ad7b3c · outbound

This paper cites and Shafer, G.

Human-Calibrated Automated Testing and Validation of Generative Language Models and Shafer, G

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:13:18.103309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.808402Z digest=sha256:e37a1b087dc4a7bfd30a94b4d9f5f0b21e9e4174034f596d26576746dc796246

Observation 834ad3c5-9147-4512-bac3-2a5cbc712681 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Human-Calibrated Automated Testing and Validation of Generative Language Models HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.813782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.813782Z digest=sha256:e067b15e86d87ee7bddd0e48acb1e4e1b6a86d750d77427d0cad52d5fbb85cdd

Observation 04f7eaa7-a087-4098-a8bc-fb02e429a0a8 · outbound

This paper cites an unresolved cited work.

Human-Calibrated Automated Testing and Validation of Generative Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:13:18.073918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T13:13:17.819874Z digest=sha256:59d1a298cfd9270e708f04a7205b9c5c0821d87fde165f422f59717f2c0bfeb5

Observation d0e46e76-711b-4f72-ad47-c0db79709e49 · outbound

This paper cites A Survey of Large Language Models.

Human-Calibrated Automated Testing and Validation of Generative Language Models A Survey of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T13:13:17.824953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:13:17.824953Z digest=sha256:4cff4bcc4e59e4d1ddf7b9d526ec89d1f0e13900321ebbd59bf5846d83e090a7

Pith citing papers

Observation 19fe34e4-4d2f-4d07-ac69-b5e58bd3892a · inbound

As It Was: Aligning LLM Search Evaluation with Historical User Preferences cites this paper.

As It Was: Aligning LLM Search Evaluation with Historical User Preferences Human-Calibrated Automated Testing and Validation of Generative Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:36:43.167106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-02T06:34:54.572363Z digest=sha256:babce8318f3ff043bda934957dc2fd1bc9f55756d37e0e5419ba7f4dfab7317c