Pith. sign in

Paper Citation Record · LEDGER

Statistical Hypothesis Testing for Auditing Robustness in Language Models

As of 8 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2506.07947.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07947 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:28:11.181361Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved15
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29d4592b-fa06-4d7c-803b-04f41a7120f1 · outbound

This paper cites Re-evaluating Evaluation in Text Summarization.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Re-evaluating Evaluation in Text Summarization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.120511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.120511Z digest=sha256:b2dd49b27e2773adcd4b2f8e52f38c5f0d709e1de0d767898b6fa7d0b08b21c0

Observation 77c6bbc9-e814-4d60-a18c-b83fa3f9270d · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Explaining and Harnessing Adversarial Examples

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.142515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.142515Z digest=sha256:2b431b6995dea6c10d88612d7ba48d142b16eb099ce62a55bde8b236460bfacf

Observation 4bbcbe76-0111-4f28-b4b2-36f41321738c · outbound

This paper cites Reducing Gender Bias in Abusive Language Detection.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Reducing Gender Bias in Abusive Language Detection

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.159775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.159775Z digest=sha256:9659e5daeeaca1b28f71749c9db9b8472994f62a3911cd76843c7de2ed6b0952

Observation 2cd5e4bc-436f-4365-9b47-c76d90088f13 · outbound

This paper cites Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Large Language Models Help Humans Verify Truthfulness -- Except When They Are Convincingly Wrong

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.165754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.165754Z digest=sha256:ea8aec31a47d8054ce343b44c60d5eea2cb0db7e071b82d9674a0e0458aefb5e

Observation 428fe31f-9dfc-4ecf-9332-74760eb580f7 · outbound

This paper cites RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models.

Statistical Hypothesis Testing for Auditing Robustness in Language Models RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.168723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.168723Z digest=sha256:978b07ece78b3080510abbb627dcc521fdf1cf0c5cd01309a043ee8220969a7a

Observation 17e6e69e-9328-4e50-8104-cb8b67cf9315 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Statistical Hypothesis Testing for Auditing Robustness in Language Models ReAct: Synergizing Reasoning and Acting in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.171948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.171948Z digest=sha256:d2548841fd02a4ef5bff38bc29938f9cdeef253f77ee39ab2275a85df9403d4f

Observation 1ae77d9e-a123-457e-97d1-4f6f0074a7b8 · outbound

This paper cites MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance.

Statistical Hypothesis Testing for Auditing Robustness in Language Models MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.177892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.177892Z digest=sha256:08cd8896612ee25870de2dbc77fd0f1c04098a1785f758697304bf2c0644cb2f

Observation 4d8d9b65-ca47-4d27-99ad-70cf5a3378af · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Statistical Hypothesis Testing for Auditing Robustness in Language Models BERTScore: Evaluating Text Generation with BERT

Reference 2003

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.174703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.174703Z digest=sha256:e2d014e1ccebad5c7c57869cdf37ece75d18abe60730d4681c985e7e18e87987

Observation b73b1547-da50-4877-909b-94f8ff8fb7fb · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Lost in the Middle: How Language Models Use Long Contexts

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.156193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.156193Z digest=sha256:7984ed096918f2f190217d88bc11f220a529af33f0e8bdd960c962d2a5be9dc0

Observation dc38017d-7887-44c6-9ea6-aedfd7cd2b39 · outbound

This paper cites Better Zero-Shot Reasoning with Role-Play Prompting.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Better Zero-Shot Reasoning with Role-Play Prompting

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.152894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.152894Z digest=sha256:977314eefb34e72beed2232c0c79ed8e01a720f2883b1587df3c3fb5d1067dd0

Observation 5d19118c-c010-4eab-a08a-2495128f71d3 · outbound

This paper cites The Effect of Sampling Temperature on Problem Solving in Large Language Models.

Statistical Hypothesis Testing for Auditing Robustness in Language Models The Effect of Sampling Temperature on Problem Solving in Large Language Models

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.162864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.162864Z digest=sha256:d8b7b2985a0892048b8898762064d8eeb33cc4cf918c36b8ba26497a80b42891

Observation 147bf186-03e4-45d1-ab88-b11ee143e6be · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Reasoning with Language Model is Planning with World Model

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.146164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.146164Z digest=sha256:5eb98f8acee44844e5efcf23367a76b42788c987eaaee33f7c0e2e32f7cdb95f

Observation e67ac1fe-e122-4f2a-bcdb-f818cd8f996f · outbound

This paper cites Accountability of AI Under the Law: The Role of Explanation.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Accountability of AI Under the Law: The Role of Explanation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.135904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.135904Z digest=sha256:41afd088947e6088448aa76b5aeccd7839f66539a06fb98254202083145e326f

Observation 85c28076-f0a8-4bda-96a6-365fa331faa0 · outbound

This paper cites The Oscars of AI Theater: A Survey on Role-Playing with Language Models.

Statistical Hypothesis Testing for Auditing Robustness in Language Models The Oscars of AI Theater: A Survey on Role-Playing with Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.128582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.128582Z digest=sha256:aa77f6e0e84a74c0ea1a44f6e5c9c982bce819dfd55db01a6a2c5860d62ca469

Observation d4e88715-6dc5-4cf4-8afd-4b96ce44151b · outbound

This paper cites Nuanced metrics for measuring unintended bias with real data for text classification.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Nuanced metrics for measuring unintended bias with real data for text classification

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:28:11.377978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:28:11.124720Z digest=sha256:ebaf603ff9d270f844b9456a076c24cd225c17aa3a4b3f9aadaf2a5fe8fe1bf8

Observation d28078df-5dbd-4f44-8ddd-96eab8434b9c · outbound

This paper cites H., and Beutel, A.

Statistical Hypothesis Testing for Auditing Robustness in Language Models H., and Beutel, A

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:28:11.359472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:28:11.139337Z digest=sha256:47078619d1a3a597e42b4286cb18c8d53c5efd07337fac2ae8bde1ef530eb02d

Observation 0f2037af-455c-43a2-ba48-929b64b276c9 · outbound

This paper cites Inherent Trade-Offs in the Fair Determination of Risk Scores.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Inherent Trade-Offs in the Fair Determination of Risk Scores

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:28:11.149472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:28:11.149472Z digest=sha256:bed29bf2249daffe5991baa8f962a64de655bf37509601bcd6312f6cab73977f

Observation cc23a211-5d07-40e3-ab79-55eafc25767c · outbound

This paper cites Measuring and mitigating unintended bias in text classification.

Statistical Hypothesis Testing for Auditing Robustness in Language Models Measuring and mitigating unintended bias in text classification

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:28:11.368946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:28:11.132613Z digest=sha256:5e16d5e8214248446680a0888488f11457cd95b9297170f6633417a03182950d

Observation 4664fcd8-01d4-434f-8881-e3dba60bafb1 · outbound

This paper cites Dollar amounts are shown in parentheses for costs exceeding 100¢ (1 USD).

Statistical Hypothesis Testing for Auditing Robustness in Language Models Dollar amounts are shown in parentheses for costs exceeding 100¢ (1 USD)

Reference 2025

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:28:11.349868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:28:11.181361Z digest=sha256:5a4eae07f704ac60c500914bdeff1beeea69bd52c0968198380cb43a61f711ae

Pith citing papers

No inbound Pith citation observations are available.