Pith. sign in

Paper Citation Record · LEDGER

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation

As of 7 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 3 inbound Pith citation observations for arXiv:2606.15127.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.15127 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T04:43:04.740925Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:01:22.131040Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T23:16:36.698169Z

Reference resolution

14 of 14 outbound references displayed

  • verified exact10
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 447b601f-3e95-4e48-9367-8fbd2feb2c6d · outbound

This paper cites Chain-of-Thought Reasoning In The Wild Is Not Always Faithful.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T16:58:43.406916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:eca23550720b7516f879c604f7753ef632bb675bb1326065d4bc8e9a0f05a4f6

Observation 60250f88-3172-43ba-bf75-dd6341c2c6f1 · outbound

This paper cites Reasoning Models Don't Always Say What They Think.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation Reasoning Models Don't Always Say What They Think

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T16:58:43.404342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:a22c5f31cb0cb3d3c4fe4d6d304cf7b4fcd09d6d5076c1c6adba9434782c99b2

Observation 0c74a992-f189-46c3-9fa2-7788131d90d3 · outbound

This paper cites Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T16:58:43.432751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:8d9ec9b34117e75914da2926e01203ae881c014eee3914a577daff4146d20078

Observation ef290be6-4202-4db5-98df-be8aaead5094 · outbound

This paper cites CURE:Circuit-Aware Unlearning for LLM-based Recommendation.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation CURE:Circuit-Aware Unlearning for LLM-based Recommendation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T16:58:43.429997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:d4234fcc123ae579eb4213f340df0b02d85407d4b67fe5bbc13bc73c8194699d

Observation 7fee1767-a400-4a05-b1de-dc869ce1bc54 · outbound

This paper cites CRAB: Codebook Rebalancing for Bias Mitigation in Generative Recommendation.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation CRAB: Codebook Rebalancing for Bias Mitigation in Generative Recommendation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T16:58:43.418529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:73852d626422db63eb6c2424c24b07c349a72f234b281cb8d4a50c60cf6ff2cf

Observation 233b53d9-2659-49c3-b08a-20235ea4514c · outbound

This paper cites M., Li, Z., Wu, X., Visweswaran, S., and Wang, Y.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation M., Li, Z., Wu, X., Visweswaran, S., and Wang, Y

Reference 6

Resolution
metadata mismatch
doi, observed 2026-06-27T04:50:34.964469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:dd4f50da86331ceccc5b236e872f1dfeb38e2a48349d27990aeb1759e34dc967

Observation 6d52b44d-ad2f-44b9-8289-36360c2938c1 · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T16:58:43.434936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:dcaaa23deae8a42c119fdf34c7916322d45dcac7a6574e2b32e01b9697708c28

Observation f66aff0b-e9f9-4bcf-93bb-4d35cf604577 · outbound

This paper cites Lin, J., Zhu, C., Kneuertz, P.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation Lin, J., Zhu, C., Kneuertz, P

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:58:43.415736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:d3925f43b23af218af85d5c4c3881576338085f5f9cc015974203bc10e631f51

Observation fc5f91cc-f1de-4cc8-9eec-600b4abba7f4 · outbound

This paper cites Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T04:50:34.962781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:95236e9741eb14b274e4f1a0df632c97d3d9113c07f1ce8fc1ba46e5cdfa1bcf

Observation a61174eb-32e5-4925-bd92-5d5fe7ea1178 · outbound

This paper cites FAIntbench: A Holistic and Precise Benchmark for Bias Evaluation in Text-to-Image Models.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation FAIntbench: A Holistic and Precise Benchmark for Bias Evaluation in Text-to-Image Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:58:43.418981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:2116dfa292b47710c673486f1ec7af284e1e23060ce48ee0b6ec90419d4ea207

Observation 9f2226b5-0f54-4c49-b3a4-2a2b8140b0df · outbound

This paper cites AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation AtelierEval: Agentic Evaluation of Humans & LLMs as Text-to-Image Prompters

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T16:58:43.413110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:13790225a2ab97bcd235c44c8031393bf0766c0a1dc0a11137d38673d1eb99f4

Observation 2207baad-4d7e-4fea-b47f-2ace7c490240 · outbound

This paper cites Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T16:58:43.427689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:6e70492df392df91e2ed83d12f61e7294c4070c20747ae9f2900d1ad521cb278

Observation 123b41f0-d534-434c-a872-d8e802ff5717 · outbound

This paper cites When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T16:58:43.426521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:167c8950d0942d98a8d823da8e77c3e94ca2a3241514ddcbec1d47f41bbd4f01

Observation 236ea5ec-cb9c-41e6-ae8a-ce87bd908d21 · outbound

This paper cites arXiv preprint arXiv:2508.15126 , year =.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation arXiv preprint arXiv:2508.15126 , year =

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:58:43.429126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:bc3818162e8d8f41f53bf2e5ff545f599e6e6dbca20b9a53e8bf1c4709b90a92

Pith citing papers

Observation 49ae5f4f-492d-4d08-86d6-394f5243dc46 · inbound

Evaluating LLM Robustness Under Domain-Specific Prompt Perturbations in Public Health Applications cites this paper.

Evaluating LLM Robustness Under Domain-Specific Prompt Perturbations in Public Health Applications Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-09T23:16:36.700008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-09T23:07:45.392658Z digest=sha256:ec6ca7690557822e6674c9839aeed919165b4946363dd2acf90d420f3e1612ed

Observation 6c8c60fb-bb4c-4cbc-bdd1-5eb025eeb94b · inbound

Phantom Guardrails: When Self-Improving Agent Harnesses Fix Failures That Never Happened cites this paper.

Phantom Guardrails: When Self-Improving Agent Harnesses Fix Failures That Never Happened Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T07:01:22.131040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:01:22.131040Z digest=sha256:d4b7cc7295529eefd5eff1e37596da24c1d3e4b5451dc42a63a036fe35888869

Observation c7e7b098-61eb-4a79-b54e-45e114912c23 · inbound

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions cites this paper.

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T15:03:42.817118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:03:42.817118Z digest=sha256:66586a57c30468ca98463bbc31159802b7bf219a391d834a1a9d6467e3fed4b9