Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:14:14.618830Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2507.18742.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:14:14.618830Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T18:50:51.463421Z
A source-named dated measurement, never combined with another source.
Source: cited_works
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 92300e3e-88d9-4da2-9150-53f00ceaa254 · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Concrete Problems in AI Safety
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a99be13-5085-41e5-b4c9-ca4d169b42a0 · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Constitutional AI: Harmlessness from AI Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d580fcf6-dbc9-40b4-99f7-9cdf28ae70b5 · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 876c917d-e201-45f0-9bcf-bc7fe7763644 · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Meta SC : Test-time safety specification optimization for language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cb297aff-7114-465e-9264-f0fe276ab9d5 · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Self-refine: Iterative refinement with self-feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd032ca6-6ea0-4d40-b09f-50bf906ddaa2 · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b2ae65ad-358f-4117-a267-578ab09b979c · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Training language models to follow instructions with human feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6a82850-c8ad-4539-8030-f72180d66015 · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6dd4117-70e2-43b0-95cb-a3c55cb0a460 · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Feedback loops with language models drive in-context reward hacking
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2461dfed-3890-4700-82f6-afee731e73bf · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9d525264-2aaa-41aa-911e-33df0121d2ab · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement A Theoretical Understanding of Self-Correction through In-context Alignment
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bdeba25-7c1a-4740-a7a3-8f016ad5ce7c · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Reward hacking in reinforcement learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e84dd880-69c4-4c75-ae14-3285da599b18 · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement write newline
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be3bb819-8425-4c9f-9981-628eba89f808 · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement @esa (Ref
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0adc396-1514-426a-a319-edb4bf246f4f · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e0ce294-199e-42bf-ba85-ded73cc29600 · outbound
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b6fb5169-8fe0-4d29-8e11-b3bf1416498c · inbound
Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.