Pith. sign in

Paper Citation Record · LEDGER

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

As of 21 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.10209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10209 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:17:59.862922Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a940407a-d1b7-4435-90f9-b189062a0605 · outbound

This paper cites Concrete Problems in AI Safety.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Concrete Problems in AI Safety

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.686636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.686636Z digest=sha256:5ef05c0863fc1e57c27477287ef1b508e311fc34663a6ad3e2d1f6debb4d1af0

Observation 1bc19727-d728-4384-be6e-5f8a6052a8a3 · outbound

This paper cites Measuring political bias in Claude.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Measuring political bias in Claude

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:01.210935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.693517Z digest=sha256:7ffa5f193252fff01cf5b4b7b085cc392954f358809adad78ce6157a77ba2bc6

Observation 3ab7e301-3aa1-4910-8a42-51fbe73230b0 · outbound

This paper cites The Internal State of an LLM Knows When It's Lying.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Internal State of an LLM Knows When It's Lying

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.698106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.698106Z digest=sha256:064caa18e53901ea06a64d50b5a1532bb743abf37ceb800c805be8b007f4afef

Observation 0728dbc1-5e0a-4ce9-82c5-b695b11f014e · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.702774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.702774Z digest=sha256:f2501d4cd5c2abc7ee835c5d402e6843c8a4126c8cd9d3f3b70bb6f0c8069874

Observation a5248c68-d469-4d80-9043-a6a7c99f8755 · outbound

This paper cites Boerner, Stephen Deems, Thomas R.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Boerner, Stephen Deems, Thomas R

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.708085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.708085Z digest=sha256:16b53062a4ca4ecf63874f166e3e2264f0869ec63ffecc3ee44a80e3fb52ad1d

Observation ad5c8121-5601-44e6-8bce-716bff1b6c15 · outbound

This paper cites Discovering Latent Knowledge in Language Models Without Supervision.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Discovering Latent Knowledge in Language Models Without Supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.712832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.712832Z digest=sha256:6ffbe7944cfd1828c45d6d8fcfd191efabd09e7a16cc402dc88fe12873a00940

Observation d314d827-13af-455e-ad6c-9ebf322be669 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.721800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.721800Z digest=sha256:e6f13871ab9aaf86ea7490ba514e0e6ada59fc7d0c7c8d245261bbcd93a18b3a

Observation 1532d86c-f36f-4370-bcec-d38badff26c4 · outbound

This paper cites Eliciting latent knowledge: How to tell if your eyes deceive you.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Eliciting latent knowledge: How to tell if your eyes deceive you

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:01.186919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.729878Z digest=sha256:936b241d6942080f40df56ecccede56be4baafa33be2224c0e642ae9486f39c0

Observation ad8e74f6-8a6a-4dca-aa03-5f6944d8fec8 · outbound

This paper cites Deep reinforcement learning from human preferences.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Deep reinforcement learning from human preferences

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.734216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.734216Z digest=sha256:f6075b17dd4f5461fbae1c297692fd19f68090b47591c58bc26993d2f8032c50

Observation c4face5c-2283-4977-9f4f-19466581e0e7 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.739636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.739636Z digest=sha256:6f95bef14482d66336977140e68f280a98c47cd35eea0df3d46632823a26d627

Observation 8deeab91-14d7-4f38-a260-c36fdde05c1d · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes QLoRA: Efficient Finetuning of Quantized LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.746751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.746751Z digest=sha256:a0178ffff032dba2629557a786f6bccc4c8b5873474386501929e1a296b64911

Observation 1660b117-16f9-414a-a927-d7ebe0524aea · outbound

This paper cites On the Relationship between Truth and Political Bias in Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes On the Relationship between Truth and Political Bias in Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.751481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.751481Z digest=sha256:3dcf2fe3f089fc2ddbfe1ff3dc0187d531aadef0a4237939a349002cd23f0317

Observation a3ef1fa3-574b-4cfa-90a8-f5f173373f36 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.756004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.756004Z digest=sha256:b73d48acfffcdf13b9f60672de1e6b4edb33ff0ba5f83f7178232de38ae847f6

Observation 4ccd811d-2e17-4c4d-a8a6-218685ff43cb · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.760327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.760327Z digest=sha256:01353ae9b0b4d193c902ff17970f3fc987536fdac7a5abeae5cf5b1250e8ba8e

Observation 07504f46-8113-4612-b4ea-b006646da988 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Language Models (Mostly) Know What They Know

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.767334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.767334Z digest=sha256:d9ad82bf268206f86fbcbd8b635d2ed2d8e794dba93cf1aaaf17a535e9f88f45

Observation 2ec66086-0e3d-43db-936c-d436265fe252 · outbound

This paper cites SGD on Neural Networks Learns Functions of Increasing Complexity.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes SGD on Neural Networks Learns Functions of Increasing Complexity

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.772617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.772617Z digest=sha256:9d288af688d1d35720780062fba8b6ac4b2f628afce05d08287f8813f4ff4b9b

Observation 2e957e4c-300d-4875-9541-0ca9b905fcd1 · outbound

This paper cites Natural emergent misalignment from reward hacking in production RL , 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Natural emergent misalignment from reward hacking in production RL , 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.777339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.777339Z digest=sha256:9ddccdbc0226236a36096f7e588cd198bcc783b99a949f6c41d2aa2c0414840d

Observation 768a663e-8244-4445-a707-0a90224fa3fe · outbound

This paper cites Categorizing Variants of Goodhart's Law.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Categorizing Variants of Goodhart's Law

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.782461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.782461Z digest=sha256:b29bd515258be6530b5fd66ecf79eaacacf9b224b0e945ab40e6f0627549999a

Observation 745f7d37-64b8-4d71-ac37-90849517e8c3 · outbound

This paper cites The Alignment Problem from a Deep Learning Perspective.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Alignment Problem from a Deep Learning Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.786686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.786686Z digest=sha256:8025163e096d0add329084c57c43a6d7046f5c888d839ed940125335b5b8e93e

Observation 2dd13208-9c3f-41db-a050-f9db41382677 · outbound

This paper cites Training language models to follow instructions with human feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Training language models to follow instructions with human feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.791545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.791545Z digest=sha256:48d5b6b4e50eeb6ee17c52d29e731ba963241aac007497b8291d55a845836e69

Observation e8f00cd7-f125-4cc5-a484-0388ae6fc566 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.796035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.796035Z digest=sha256:134b79360efc596e4e6d24250e45c69368959112428f5f6ae8be510e73e2b220

Observation 484ec70b-f921-4e9e-983b-99a690623927 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Discovering Language Model Behaviors with Model-Written Evaluations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.800654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.800654Z digest=sha256:91cb1e24955e55c9d61331f5ad32a9a353007b5203f9d5cf280812aa46ed9679

Observation e1cec0d2-b5a3-4ae4-8ed2-1b58c08b1be4 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.805283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.805283Z digest=sha256:7a5ffe171eb0ae4e2f34e5cc91041a9d7b095a1b331ad55eb469bfcb5bb0d989

Observation 6124cd75-fc2f-4c8a-8dd5-1d484253f19c · outbound

This paper cites On the Spectral Bias of Neural Networks.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes On the Spectral Bias of Neural Networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.809491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.809491Z digest=sha256:e293836091594f38a645b987403c441202b8fbdb44e5694ee6d4676e30241c38

Observation 67ff7e07-67ae-45d8-8463-9eea83e2a93e · outbound

This paper cites an unresolved cited work.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:18:01.163686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.814145Z digest=sha256:8c63cca18d55dcff6a865b4cf9eb006fc53b0eca1a34a2b1e77c8713c04e937d

Observation 0f81ba7d-5427-4d0d-ab0d-64da85f6dcc6 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.822017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.822017Z digest=sha256:3b0638f6c22060346e27a39a2c96cf9461ad4067a517e3f4735f20cafdeb74a4

Observation 70b8a9c3-467a-4eff-a4ce-6e9bf5c5cdce · outbound

This paper cites Defining and Characterizing Reward Hacking.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Defining and Characterizing Reward Hacking

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.828256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.828256Z digest=sha256:9222bb15b113daf0882c597cbd0b58f41b850483728b38ff2e0ce3a1da014636

Observation 8897593c-21db-4e5e-82de-8e61f6e842a1 · outbound

This paper cites Learning to summarize from human feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Learning to summarize from human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.845655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.845655Z digest=sha256:ff7439e43bb0d66d0bbf884d8248c5787ede348eb16c2e73d2930644ea81270f

Observation ff6f8909-8f9b-422f-a18f-1469f7c91ff8 · outbound

This paper cites Inoculation prompting: Eliciting traits from LLMs during training can suppress them at test-time, 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Inoculation prompting: Eliciting traits from LLMs during training can suppress them at test-time, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.850544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.850544Z digest=sha256:1d293356d7c739f2441cf177ec5d24c2a4c2ae3db2981cca41cfeb07732b5d78

Observation 5c435b4a-93a8-44f1-a40f-6e524e070d3c · outbound

This paper cites Deep learning generalizes because the parameter-function map is biased towards simple functions.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Deep learning generalizes because the parameter-function map is biased towards simple functions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.855765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.855765Z digest=sha256:9e4594b30b0358906c46cef83fcd4a5e7e32d4de401026f40e86981e76db12e1

Observation 4f4f9820-c79c-477a-9f39-3968f834b1a3 · outbound

This paper cites Inoculation prompting: Instructing LLMs to misbehave at train-time improves test-time alignment, 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Inoculation prompting: Instructing LLMs to misbehave at train-time improves test-time alignment, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.862922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.862922Z digest=sha256:58846093641ab6b8db90cf7974e7ba51fae19a89fa9481ed7d6e5080ac669bb9

Pith citing papers

No inbound Pith citation observations are available.