Pith. sign in

Paper Citation Record · LEDGER

Beyond Accuracy: Behavioral Testing of NLP models with CheckList

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2005.04118.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2005.04118 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:39:11.047720Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T02:49:24.847192Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5caf5a5a-1988-475f-8864-1b178e1e3610 · inbound

Jailbreaking Black Box Large Language Models in Twenty Queries cites this paper.

Jailbreaking Black Box Large Language Models in Twenty Queries Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:48:33.255295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T09:48:31.721745Z digest=sha256:a421ee4f17cf3ccd74321b721c13947c352cfa5b9fe13e7d4116eab39190f263

Observation 2aac01c0-4322-45bb-8127-41ef71993d90 · inbound

The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection cites this paper.

The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:14.503149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:14.503149Z digest=sha256:8a17a6c8da082d019f2162e4a2482c30e36ac747b45a1968f5041cd989bf1947

Observation 841eeaf4-f3c7-463f-baf0-f5946d23113c · inbound

Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection cites this paper.

Enhancing Zero-shot Chain of Thought Prompting via Uncertainty-Guided Strategy Selection Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T05:33:56.155890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:33:56.155890Z digest=sha256:2c60947f8884c788c1985c9e089dac26ed95151476fb0d8eedceacd002fcdf26

Observation 394cdacf-fc77-4028-97b1-a22bfe7856bf · inbound

Counterfactual Samples Constructing and Training for Commonsense Statements Estimation cites this paper.

Counterfactual Samples Constructing and Training for Commonsense Statements Estimation Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T23:23:25.493077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:23:25.493077Z digest=sha256:bd27dda48c12b8422b1dc4f88885239ddcb6d689000c792114c079ecba918089

Observation c6fbf12b-d000-483f-9986-45df04151659 · inbound

Measuring Diversity in Synthetic Datasets cites this paper.

Measuring Diversity in Synthetic Datasets Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:50.853489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:54:50.853489Z digest=sha256:61e8c7a01dfe19cc70e05a026f1ba692b46fd2de40a22bfd8dfb3614126d5ad6

Observation 53292430-4792-4d4e-9df2-2482e8da7b43 · inbound

aiXamine: Simplified LLM Safety and Security cites this paper.

aiXamine: Simplified LLM Safety and Security Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-16T11:39:11.047720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:39:11.047720Z digest=sha256:85a15947c0d21d70549db45f3c9fb1cb250939473338d6fd8254875f6b1501a1

Observation b2342e17-2f67-40ca-8b25-6fd197c436ba · inbound

Ensuring Reproducibility in Generative AI Systems for General Use Cases: A Framework for Regression Testing and Open Datasets cites this paper.

Ensuring Reproducibility in Generative AI Systems for General Use Cases: A Framework for Regression Testing and Open Datasets Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:27:52.927658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:27:52.927658Z digest=sha256:aed88d1f997bd1b6a24c600c8d3cebc0a59d4ae3299dd2ad2d1b607d8104e584

Observation f1b3c440-7823-4e4c-b899-9957d5f700d2 · inbound

Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models cites this paper.

Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:10.675956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:14:10.675956Z digest=sha256:ad0bb669b3b07299b55bc5c5bf22f6802b05f9eb77a4c762b4e15b9470ea60e7

Observation 44185dd5-4404-45c2-b8e8-5e3dcf2f7052 · inbound

GenFair: Systematic Test Generation for Fairness Fault Detection in Large Language Models cites this paper.

GenFair: Systematic Test Generation for Fairness Fault Detection in Large Language Models Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:42.214448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:42.214448Z digest=sha256:9728911017539cca6c4d545ad4a8895df280e27b326522846fffc180d9a8bb4a

Observation 1bbb4541-0b04-4f46-83a6-dcf80abbd733 · inbound

Potemkin Understanding in Large Language Models cites this paper.

Potemkin Understanding in Large Language Models Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.993054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.993054Z digest=sha256:736f2ca609a9f75ffe452258c7f874b98bd4ce468c7f2717f108657d54226186

Observation ff9351b1-e60b-4ae7-a2cd-11fd9d455bf6 · inbound

ASSURE: Metamorphic Testing for AI-powered Browser Extensions cites this paper.

ASSURE: Metamorphic Testing for AI-powered Browser Extensions Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:43:29.244466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:43:29.244466Z digest=sha256:253a66b56e379ddcd2e0946d9ea8edd142047fc94110db936a2d5ef15d4b72e3

Observation d9db4d8b-1f4d-4a14-a274-8c4700c34de4 · inbound

Agentic Web: Weaving the Next Web with AI Agents cites this paper.

Agentic Web: Weaving the Next Web with AI Agents Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 186

Resolution
unresolved
no resolver link, observed 2026-08-06T13:05:39.875211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:05:39.875211Z digest=sha256:c1fa0c3dbb357703703d19a531ba731cdc3de7fec3eb24b2f90e9e565245c582

Observation ef847eb1-bc82-4595-8096-466b8c60d63f · inbound

Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation cites this paper.

Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T22:50:43.780478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:50:43.780478Z digest=sha256:0c628b0be4881c5518acc9a6e661d8de1460a1c25b986b1d2ab2d8fe02e216e5

Observation 997716ef-ddec-459c-8def-2bc06e2f7fb0 · inbound

SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation cites this paper.

SPARSE Data, Rich Results: Few-Shot Semi-Supervised Learning via Class-Conditioned Image Translation Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T22:46:44.391762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:46:44.391762Z digest=sha256:5fb822105f670681957291c9f5f0d507257ec11e474aac8a890bbb473ae102f0

Observation f4c8d66b-1d0a-43ca-ab53-e87408ad5479 · inbound

Evalet: Evaluating Large Language Models through Functional Fragmentation cites this paper.

Evalet: Evaluating Large Language Models through Functional Fragmentation Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:01:39.982194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T16:57:25.259866Z digest=sha256:f74a019f301974ba13638817dab611f3517f6ca0fa1ba309ba1f7371938e8bfa

Observation 8c9f4954-7740-491d-b601-f9c219ce2473 · inbound

When to Trust the Answer: Question-Aligned Semantic Nearest Neighbor Entropy for Safer Surgical VQA cites this paper.

When to Trust the Answer: Question-Aligned Semantic Nearest Neighbor Entropy for Safer Surgical VQA Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:22:15.719282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T01:21:26.854636Z digest=sha256:7427c38c857d6f1a764ef5bcd00f282ef9d4ac90bfa6c2efd9813b8cf711fcfd

Observation 26e056f2-14f0-433a-a8d8-32a364ff2adf · inbound

Learning to Discover at Test Time cites this paper.

Learning to Discover at Test Time Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:16:04.194337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T05:16:04.001700Z digest=sha256:7ffa7f8b4b88ce7ecccf442f64614296dffa0457dda59b1680d78255d8d24da3

Observation 97b09954-a9ce-4b3f-864e-8c7ddecbcdf8 · inbound

Measuring Representation Robustness in Large Language Models for Geometry cites this paper.

Measuring Representation Robustness in Large Language Models for Geometry Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:38:10.520784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T19:35:32.531660Z digest=sha256:a4ebfd18e25b86b35d8e1ad2e6363247e0b416049f309a18da37abebfb150e74

Observation 38f3e839-bca3-4ff1-a7b5-be4ac92820fc · inbound

Social Bias in LLM-Generated Code: Benchmark and Mitigation cites this paper.

Social Bias in LLM-Generated Code: Benchmark and Mitigation Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 153

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:36:08.937852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-09T19:34:51.433422Z digest=sha256:e366c6d8e0931751b262f084eaaab55e58b665399014be48c3e63c65f60b7b93

Observation 633d3a0d-5872-477f-a314-cacd2f379e40 · inbound

How to Interpret Agent Behavior cites this paper.

How to Interpret Agent Behavior Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:27:35.960649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T18:23:25.269217Z digest=sha256:c664077a65dbde2b09a7ffb0dcbc59e19ea09f0e760f9c609f3eadad6d52de70

Observation 93a4eb04-fcab-4d4a-8a79-1ffdac38e016 · inbound

Interactive Evaluation Requires a Design Science cites this paper.

Interactive Evaluation Requires a Design Science Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:58:14.003198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T10:55:08.135630Z digest=sha256:790feca4774d64e118a5ebbdd1a2cc2c04364ecb952be4eb90e2b395947ef5cc

Observation f9bb0636-5f55-41d3-b279-05e048ec2fce · inbound

Trustworthy Recommendation in the Era of Large Language Models: Opportunities and Challenges cites this paper.

Trustworthy Recommendation in the Era of Large Language Models: Opportunities and Challenges Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 195

Resolution
verified exact
arxiv_id, observed 2026-06-28T18:42:29.189519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T18:41:06.636352Z digest=sha256:d8bc674fe70fc4607881a5788411451e2a264c130d59944d1d03ae149218fa89

Observation c6b290ae-5898-429c-baf2-7260f431fd4e · inbound

CRAFT: Cost-aware Refinement And Front-aware Tuning of Prompts cites this paper.

CRAFT: Cost-aware Refinement And Front-aware Tuning of Prompts Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:46:46.394764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T06:39:17.268337Z digest=sha256:85e34bfdd242b06ac754c9ac0be48b2800f77b53aff388b2d643cc07b6cb4b85

Observation 1d509cc0-ac8e-4408-8151-33cc6236c888 · inbound

What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks cites this paper.

What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:47:31.706084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T16:18:01.850874Z digest=sha256:6daca55fd9176d1f87ff302b2f95ee68e7731455331371207b3345f72530cd36

Observation 1238b618-2488-4f13-bd1e-02eb26ae55dd · inbound

Iterating Toward Better Search: A Two-Agent Simulation Framework for Evaluating Agentic Search Architectures in E-Commerce cites this paper.

Iterating Toward Better Search: A Two-Agent Simulation Framework for Evaluating Agentic Search Architectures in E-Commerce Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:18:23.057855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T07:05:31.737761Z digest=sha256:842fc562b9e7bc218494a2d1009d79ef07eed619e6033cc6b7d7370f36627043

Observation 142e94c2-a96f-4559-95d7-a7ec100e14a1 · inbound

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning cites this paper.

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.849283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T19:04:45.062426Z digest=sha256:eb413940758790f5130c5ad7e6aad30651c58d6d0da3dddfc4b9648d3153f9c5

Observation 2e78ebc0-63ab-4619-8c6f-7b5caf2a74da · inbound

Carolina Guide: A Multi-Agent RAG System with Institutional Guardrails for Academic Policy Assistance cites this paper.

Carolina Guide: A Multi-Agent RAG System with Institutional Guardrails for Academic Policy Assistance Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:14:37.507549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T11:13:23.740219Z digest=sha256:852831d87f7af465ec88a88464b032c27e2e1dc73b15e3abda6590e776c09aab

Observation 7c711e76-babd-4a98-80a1-2ccccfe33175 · inbound

Testing Retrieval-Augmented Generation Systems with Chunk Coverage cites this paper.

Testing Retrieval-Augmented Generation Systems with Chunk Coverage Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T15:54:00.365263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:54:00.365263Z digest=sha256:0306a54951447db650918f343652b1fccd4b1f2c9fd364d38d22ba0b27eda91c

Observation dd85f4fe-5239-4949-89ba-8c79e8ae304e · inbound

Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development cites this paper.

Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T07:40:17.900025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:40:17.900025Z digest=sha256:2cddf4cf2139620ab5e7694152040dabbdf8b389d879543f5d585a738f3e1f14

Observation 48282414-8bb1-4cee-aeb7-6aafdb4a4c1f · inbound

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy cites this paper.

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy Beyond Accuracy: Behavioral Testing of NLP models with CheckList

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T13:46:09.525133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:46:09.525133Z digest=sha256:2bf9c0aafaa956d380971f9e762ee979e9eb23f135f7437c062f27cdef60241b