Pith. sign in

Paper Citation Record · LEDGER

Statistical Hypothesis Testing for Auditing Robustness in Language Models

As of 8 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2506.07947.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07947 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:28:11.181361Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved15
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29d4592b-fa06-4d7c-803b-04f41a7120f1 · outbound

This paper cites Re-evaluating Evaluation in Text Summarization.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Re-evaluating Evaluation in Text Summarization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.120511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.120511Z digest=sha256:10eaf39647102785fe51ee85e76cec5961a96aa3daa565333e81477809b26cee

Observation 77c6bbc9-e814-4d60-a18c-b83fa3f9270d · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Explaining and Harnessing Adversarial Examples

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.142515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.142515Z digest=sha256:4f4ab65d81e64f1a33fe89a943aa7676af064c5f6c4382c80f8baedd29d05f0d

Observation 4bbcbe76-0111-4f28-b4b2-36f41321738c · outbound

This paper cites Reducing Gender Bias in Abusive Language Detection.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Reducing Gender Bias in Abusive Language Detection

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.159775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.159775Z digest=sha256:a076934def54212d3c619e12de1b4fae51fa7ff02c7eee9d3dcb5211ba530dfd

Observation 2cd5e4bc-436f-4365-9b47-c76d90088f13 · outbound

This paper cites Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.165754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.165754Z digest=sha256:6c5e7ae3d238e58c3e55040c6c52b10dee75115a27d8dc97018f6a2a45105b03

Observation 428fe31f-9dfc-4ecf-9332-74760eb580f7 · outbound

This paper cites RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models.

Statistical Hypothesis Testing for Auditing Robustness in Language Models RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.168723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.168723Z digest=sha256:609722a9427a173ece4c519162272182e14e9b52c8a8295eec20904a83d3cf1d

Observation 17e6e69e-9328-4e50-8104-cb8b67cf9315 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Statistical Hypothesis Testing for Auditing Robustness in Language Models ReAct: Synergizing Reasoning and Acting in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.171948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.171948Z digest=sha256:0a71b35d403cb65f9448d7ea93f0d4e2b5984a416075647b8c2b02469609b0ff

Observation 1ae77d9e-a123-457e-97d1-4f6f0074a7b8 · outbound

This paper cites MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance.

Statistical Hypothesis Testing for Auditing Robustness in Language Models MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.177892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.177892Z digest=sha256:0eb521ff58acf88eb7ceab61c23e77da14722b85aa7049e114693c789e430637

Observation 4d8d9b65-ca47-4d27-99ad-70cf5a3378af · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Statistical Hypothesis Testing for Auditing Robustness in Language Models BERTScore: Evaluating Text Generation with BERT

Reference 2003

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.174703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.174703Z digest=sha256:17141bbb221ce1202302e10d660e2ad6ac8968767571566674da9fe6476c363c

Observation b73b1547-da50-4877-909b-94f8ff8fb7fb · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Lost in the Middle: How Language Models Use Long Contexts

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.156193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.156193Z digest=sha256:ed544a0af84c807534151fbee75ba54351c662773c5f1400c8c0a91daf31c81f

Observation dc38017d-7887-44c6-9ea6-aedfd7cd2b39 · outbound

This paper cites Better Zero-Shot Reasoning with Role-Play Prompting.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Better Zero-Shot Reasoning with Role-Play Prompting

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.152894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.152894Z digest=sha256:3672d03707bc61c80d30c432f71931b2a78cba89a576c803d75a6da8940d485c

Observation 5d19118c-c010-4eab-a08a-2495128f71d3 · outbound

This paper cites The Effect of Sampling Temperature on Problem Solving in Large Language Models.

Statistical Hypothesis Testing for Auditing Robustness in Language Models The Effect of Sampling Temperature on Problem Solving in Large Language Models

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.162864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.162864Z digest=sha256:f24bc0f5b6a83fcc0b6b47242251fa2099deed1a4bc9ad210dc5a6e5057cf252

Observation 147bf186-03e4-45d1-ab88-b11ee143e6be · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Reasoning with Language Model is Planning with World Model

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.146164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.146164Z digest=sha256:27d705f998398fd38949738c9585f73b3a01bceac2fc1656b9c38ef7f3bb916b

Observation e67ac1fe-e122-4f2a-bcdb-f818cd8f996f · outbound

This paper cites Accountability of AI Under the Law: The Role of Explanation.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Accountability of AI Under the Law: The Role of Explanation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.135904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.135904Z digest=sha256:3cad993a5088ca461e9ac0e6e6ba73816e108f9d629e9650cffa47c98d5dd748

Observation 85c28076-f0a8-4bda-96a6-365fa331faa0 · outbound

This paper cites The Oscars of AI Theater: A Survey on Role-Playing with Language Models.

Statistical Hypothesis Testing for Auditing Robustness in Language Models The Oscars of AI Theater: A Survey on Role-Playing with Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.128582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.128582Z digest=sha256:725eef63670cbe568f99e21954c334ae74baa83d9164127d142e5a169de10f64

Observation d4e88715-6dc5-4cf4-8afd-4b96ce44151b · outbound

This paper cites Nuanced metrics for measuring unintended bias with real data for text classification.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Nuanced metrics for measuring unintended bias with real data for text classification

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:28:11.377978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:28:11.124720Z digest=sha256:b2282d6357d244a34481511ccf9fb25874802dcd9ee8127a734516cb854f2070

Observation d28078df-5dbd-4f44-8ddd-96eab8434b9c · outbound

This paper cites H., and Beutel, A.

Statistical Hypothesis Testing for Auditing Robustness in Language Models H., and Beutel, A

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:28:11.359472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:28:11.139337Z digest=sha256:a631766e138b1e5440c3fb4ba4522423a7e9574abee2969a85072bf463fa191b

Observation 0f2037af-455c-43a2-ba48-929b64b276c9 · outbound

This paper cites Inherent Trade-Offs in the Fair Determination of Risk Scores.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Inherent Trade-Offs in the Fair Determination of Risk Scores

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.149472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.149472Z digest=sha256:f4a6ae3384d7ff592ce66f059ce9797069f135959d9d111c84aa81aec14d577a

Observation cc23a211-5d07-40e3-ab79-55eafc25767c · outbound

This paper cites Measuring and mitigating unintended bias in text classification.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Measuring and mitigating unintended bias in text classification

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:28:11.368946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:28:11.132613Z digest=sha256:6b8c9a01e6732010b1c96e5a02f3627222dff5496a656edac1186f5ccc65183e

Observation 4664fcd8-01d4-434f-8881-e3dba60bafb1 · outbound

This paper cites Dollar amounts are shown in parentheses for costs exceeding 100¢ (1 USD).

Statistical Hypothesis Testing for Auditing Robustness in Language Models Dollar amounts are shown in parentheses for costs exceeding 100¢ (1 USD)

Reference 2025

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:28:11.349868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:28:11.181361Z digest=sha256:ea74eb9407bda5fc243eb551fd752eb65976baebf0275f9db8273c531c495f5a

Pith citing papers

No inbound Pith citation observations are available.