Pith. sign in

Paper Citation Record · LEDGER

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement

As of 21 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2507.18742.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18742 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:14:14.618830Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T18:50:51.463421Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 92300e3e-88d9-4da2-9150-53f00ceaa254 · outbound

This paper cites Concrete Problems in AI Safety.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Concrete Problems in AI Safety

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T18:14:14.543125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:14:14.543125Z digest=sha256:8103175fdcfd8c5c96aa92eca1ae9bbed767eb259bdd7e1141cc8d6a68def831

Observation 7a99be13-5085-41e5-b4c9-ca4d169b42a0 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Constitutional AI: Harmlessness from AI Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:14:14.548634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:14:14.548634Z digest=sha256:3229e25cccd632851742ca04f904ef7982039d95b24dd61b922b01edd5b1121d

Observation d580fcf6-dbc9-40b4-99f7-9cdf28ae70b5 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T18:14:14.553499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:14:14.553499Z digest=sha256:8f42298e2e69b2d96050dfcab1ecadf167e8ed40b8291e08ad89229a57b0093c

Observation 876c917d-e201-45f0-9bcf-bc7fe7763644 · outbound

This paper cites Meta SC : Test-time safety specification optimization for language models.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Meta SC : Test-time safety specification optimization for language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:14:14.844820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:14:14.558346Z digest=sha256:9ba684092b3d2c8356f539b7d1927d453d15686e249e4d0dd45874e1bd4a8876

Observation cb297aff-7114-465e-9264-f0fe276ab9d5 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Self-refine: Iterative refinement with self-feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:14:14.563129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:14:14.563129Z digest=sha256:4929f7cc511033818873b95752c3fba2faf5db51b10c6d785b067c5f3bf6f585

Observation bd032ca6-6ea0-4d40-b09f-50bf906ddaa2 · outbound

This paper cites Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:14:14.688788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:14:14.568282Z digest=sha256:340837371b3712d1a532cbd9fbd78e368974aaa5c5a4a7377ff186906d4f72db

Observation b2ae65ad-358f-4117-a267-578ab09b979c · outbound

This paper cites Training language models to follow instructions with human feedback.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Training language models to follow instructions with human feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:14:14.573132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:14:14.573132Z digest=sha256:4397752df47227f9693f16160b9c4b46777948c0a6c80f5b4e416a9fe63255d6

Observation e6a82850-c8ad-4539-8030-f72180d66015 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:14:14.578328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:14:14.578328Z digest=sha256:983ad32c5e88828359a3a5a64b0f5a7c25569a65abf9a56855c70976f96d352f

Observation d6dd4117-70e2-43b0-95cb-a3c55cb0a460 · outbound

This paper cites Feedback loops with language models drive in-context reward hacking.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Feedback loops with language models drive in-context reward hacking

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:14:14.812064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:14:14.583011Z digest=sha256:27d6e240f2075e78d1c1be1c23ab660d8e32f2c8c823f6a18918110cc66f40b8

Observation 2461dfed-3890-4700-82f6-afee731e73bf · outbound

This paper cites an unresolved cited work.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:14:14.796642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:14:14.587262Z digest=sha256:dd88331be8628fcd39fd5c6172799703b8a566e65c5d445bf3769601bf2aa094

Observation 9d525264-2aaa-41aa-911e-33df0121d2ab · outbound

This paper cites A Theoretical Understanding of Self-Correction through In-context Alignment.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement A Theoretical Understanding of Self-Correction through In-context Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:14:14.591491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:14:14.591491Z digest=sha256:505af2db3116c883fde81131a9d64ca38b8c044bd47ea1baa1a016a5795b1d4d

Observation 7bdeba25-7c1a-4740-a7a3-8f016ad5ce7c · outbound

This paper cites Reward hacking in reinforcement learning.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Reward hacking in reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:14:14.596987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:14:14.596987Z digest=sha256:b663700dd59969a9dc96362d8be890642fcb9424f0c995016ca866d0b6cf579b

Observation e84dd880-69c4-4c75-ae14-3285da599b18 · outbound

This paper cites write newline.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement write newline

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:14:14.603803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:14:14.603803Z digest=sha256:21a9bd3a504cbfbe6d8a99284854f0c6f76229512638200cb0a40380b7782f20

Observation be3bb819-8425-4c9f-9981-628eba89f808 · outbound

This paper cites @esa (Ref.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement @esa (Ref

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:14:14.609788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:14:14.609788Z digest=sha256:7e68b2dc1fc841224d7cacfa0fb87cfec1b58f772ef3dcbad3a1b46926b875ba

Observation d0adc396-1514-426a-a319-edb4bf246f4f · outbound

This paper cites an unresolved cited work.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T18:14:14.614450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:14:14.614450Z digest=sha256:ae3e8ebbcfe00647640d24a5b72ebcc0ed8660fa90b925c2c7c7b6e3e9885b7e

Observation 2e0ce294-199e-42bf-ba85-ded73cc29600 · outbound

This paper cites an unresolved cited work.

Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:14:14.745525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:14:14.618830Z digest=sha256:8b3fcb422995076e9070993c26cff63b478269a492b4dad772b0d1a8b28443e9

Pith citing papers

Observation b6fb5169-8fe0-4d29-8e11-b3bf1416498c · inbound

Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents cites this paper.

Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T18:50:51.463421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T18:50:51.463421Z digest=sha256:5dff2efb8a59293570f145dbeeaee1fbb643014f79b3a8cf653f0b3e0aeaa353