Pith. sign in

Paper Citation Record · LEDGER

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop

As of 22 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2607.23002.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23002 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:57:40.804853Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c72079c3-cb88-4ff8-bbd8-11d56d9f99b7 · outbound

This paper cites HITS: High-coverage LLM-based Unit Test Generation via Method Slicing.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop HITS: High-coverage LLM-based Unit Test Generation via Method Slicing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.762460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.762460Z digest=sha256:ede643fba1cf77eca181717cfeaa9b600da9178e637f74cc15150ead57be8149

Observation 77756427-eb8a-4b8b-b500-35275fd6bb24 · outbound

This paper cites LLM Evaluators Recognize and Favor Their Own Generations.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop LLM Evaluators Recognize and Favor Their Own Generations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.766985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.766985Z digest=sha256:2d8cd0849346b7b99523a1f6d78fa1c79cb96231bdc18a7af3354597d68ff7b2

Observation d1890c78-32a4-43c5-9ac1-cc910ac0fc50 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.771519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.771519Z digest=sha256:62c416d939aaed3e73a68344e52566d62b0abb22b89e79d4cd09b56d00910136

Observation 5edd23c8-e32f-443d-abcc-7c43b8f7eec3 · outbound

This paper cites Great Models Think Alike and this Undermines AI Oversight.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Great Models Think Alike and this Undermines AI Oversight

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.784183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.784183Z digest=sha256:c9f9ab177b59550e1611f65436c2130d39c3b168e3e6be306ef7523e19cbac47

Observation b96d9aec-229d-49c0-8c0f-0037c26bb2c9 · outbound

This paper cites TestForge: Feedback-Driven, Agentic Test Suite Generation.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop TestForge: Feedback-Driven, Agentic Test Suite Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.788358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.788358Z digest=sha256:5326696940581374998a3915d77aba1e3dd94f1969cd1c19ed100fbc73204379

Observation ec47ea65-66f6-4941-ac0c-7d60aa9fefde · outbound

This paper cites Registered Reports in Software Engineering.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Registered Reports in Software Engineering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.792701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.792701Z digest=sha256:403839996cec6ead98f807186f595ea796f893f15d2e8d76ce87cc1b306bd482

Observation 22056d44-b81a-4ae6-bab1-f2b8f129ee90 · outbound

This paper cites Fuzz4All: Universal Fuzzing with Large Language Models.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Fuzz4All: Universal Fuzzing with Large Language Models

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.804853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.804853Z digest=sha256:a7e85424e775ecd94393ec948f2b8bc6993c7320724ebf56b1b3745ad6938463

Observation 0b038097-1569-4f96-8f5a-a1e3dc6c6fdf · outbound

This paper cites Mutation Testing Advances: An Analysis and Survey.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Mutation Testing Advances: An Analysis and Survey

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.796638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.796638Z digest=sha256:6f494909d03dee47cb7de7f41bc11ade0dc02586d0e29580ee7810dba7ffd026

Observation 492a5aa9-ba2e-4100-ae62-220eb02b0a47 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.800920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.800920Z digest=sha256:a72d8d743ffae19a7ab4c72baea8ed39581b5775c2f0c9ac3dee626462c3ea74

Observation 9d0c53ab-c15b-4822-816e-caade29f8723 · outbound

This paper cites Pynguin: Automated Unit Test Generation for Python.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Pynguin: Automated Unit Test Generation for Python

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.758128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.758128Z digest=sha256:ae9bb3824c42df779cab96ae78a063ff91c6be6b458a2481479ff66864db8186

Observation 693eb81b-ba9b-48e2-aa95-db5bfc04361a · outbound

This paper cites An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.753104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.753104Z digest=sha256:146367bd9225b9095a843f323578066757bd2dca9819a37836053c9c42f6919f

Observation ba55738c-e6e6-46c6-b77c-84eb57d5b856 · outbound

This paper cites Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.775727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.775727Z digest=sha256:2fb38fea18c61679c0e0f3ddee60059bc6a380ab56fa4fde7d5d822c01a92df0

Observation e272ef24-f536-4312-a1da-6ec30669cf8b · outbound

This paper cites Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery.

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T03:57:40.779886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:57:40.779886Z digest=sha256:d3ef509fbe71c8b576f50cf08a642aa235f831d92af349a6b08561361c3122a8

Pith citing papers

No inbound Pith citation observations are available.