Pith. sign in

Paper Citation Record · LEDGER

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

As of 14 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.10209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10209 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:17:59.862922Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a940407a-d1b7-4435-90f9-b189062a0605 · outbound

This paper cites Concrete Problems in AI Safety.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Concrete Problems in AI Safety

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.686636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.686636Z digest=sha256:6404d2827f4510240c7237355c7c8877ecc2c3c412783eafea149f3b289eaed3

Observation 1bc19727-d728-4384-be6e-5f8a6052a8a3 · outbound

This paper cites Measuring political bias in Claude.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Measuring political bias in Claude

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:01.210935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.693517Z digest=sha256:309209f5f398c42180c19ed24bd85e967101d85b995592ae66a379a55305c762

Observation 3ab7e301-3aa1-4910-8a42-51fbe73230b0 · outbound

This paper cites The Internal State of an LLM Knows When It's Lying.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Internal State of an LLM Knows When It's Lying

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.698106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.698106Z digest=sha256:572d171140c33ae2673062a70224c8b7f604bfbb3630b2118676d553a864c03b

Observation 0728dbc1-5e0a-4ce9-82c5-b695b11f014e · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.702774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.702774Z digest=sha256:275c7dec45c20b85107f2a7c5a0c3d0b7c6b1f4f6c6be2b4beacbbc83ad112b2

Observation a5248c68-d469-4d80-9043-a6a7c99f8755 · outbound

This paper cites Boerner, Stephen Deems, Thomas R.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Boerner, Stephen Deems, Thomas R

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.708085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.708085Z digest=sha256:32b623d7c575edeee95325be9805dbcd6cd2874bf3a31bc2015d8b99cc1fcbcb

Observation ad5c8121-5601-44e6-8bce-716bff1b6c15 · outbound

This paper cites Discovering Latent Knowledge in Language Models Without Supervision.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Discovering Latent Knowledge in Language Models Without Supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.712832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.712832Z digest=sha256:5b3ce2ba86c4cf532a626ebf41f4264c0e5a776d48eb4c7f18b6d4317bf7a7d7

Observation d314d827-13af-455e-ad6c-9ebf322be669 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.721800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.721800Z digest=sha256:5340a7068896cad12570d74f17490c4fde46aecc65b047de3c2d326041032a91

Observation 1532d86c-f36f-4370-bcec-d38badff26c4 · outbound

This paper cites Eliciting latent knowledge: How to tell if your eyes deceive you.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Eliciting latent knowledge: How to tell if your eyes deceive you

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:01.186919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.729878Z digest=sha256:f91c5650c0eef5f847a917512dbecaf1bb9b6e062914f64d0a9871b87739e1e7

Observation ad8e74f6-8a6a-4dca-aa03-5f6944d8fec8 · outbound

This paper cites Deep reinforcement learning from human preferences.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Deep reinforcement learning from human preferences

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.734216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.734216Z digest=sha256:f646ba78df459ccaf24f4664ed5edc659b3bf591072a8f9bc1053353bcbf344a

Observation c4face5c-2283-4977-9f4f-19466581e0e7 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.739636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.739636Z digest=sha256:10237069b43a95ef194e4e6f907937150ebf120e959ded6b595d7275958b4bda

Observation 8deeab91-14d7-4f38-a260-c36fdde05c1d · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes QLoRA: Efficient Finetuning of Quantized LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.746751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.746751Z digest=sha256:904e2d4204852d5b6b122f804815a32c8cac508c71c1d19cad29a2cd0312a898

Observation 1660b117-16f9-414a-a927-d7ebe0524aea · outbound

This paper cites On the Relationship between Truth and Political Bias in Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes On the Relationship between Truth and Political Bias in Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.751481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.751481Z digest=sha256:10880392782e256d06b0cee133633b5c2878c79c9daa1e8a579c9f1a4e301271

Observation a3ef1fa3-574b-4cfa-90a8-f5f173373f36 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.756004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.756004Z digest=sha256:b464f8d7158fa5e4c86f9c459255f6b7e2aaeb3eee333c6ebee74bbd41ef1767

Observation 4ccd811d-2e17-4c4d-a8a6-218685ff43cb · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.760327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.760327Z digest=sha256:65d05c4e409fd5884ac1274817377019c456d568626e0aba484b852020dd4b9d

Observation 07504f46-8113-4612-b4ea-b006646da988 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Language Models (Mostly) Know What They Know

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.767334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.767334Z digest=sha256:a21aee17b0d7e246ea458af8f58d80b7f667c532a648ea0e7dc340cd06e2ee28

Observation 2ec66086-0e3d-43db-936c-d436265fe252 · outbound

This paper cites SGD on Neural Networks Learns Functions of Increasing Complexity.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes SGD on Neural Networks Learns Functions of Increasing Complexity

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.772617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.772617Z digest=sha256:89f00ab84c8c75cae3cf361725a589b47da88af3d850a899721256cacff9fb6f

Observation 2e957e4c-300d-4875-9541-0ca9b905fcd1 · outbound

This paper cites Natural emergent misalignment from reward hacking in production RL , 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Natural emergent misalignment from reward hacking in production RL , 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.777339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.777339Z digest=sha256:e0310775c9ad50380c71f104c69a1e10e48e9df47ab6a2254b213d351c195173

Observation 768a663e-8244-4445-a707-0a90224fa3fe · outbound

This paper cites Categorizing Variants of Goodhart's Law.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Categorizing Variants of Goodhart's Law

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.782461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.782461Z digest=sha256:24b1516181d5b9f47c0dc650174bc16943d30f163769ea02e975ad68aa310c49

Observation 745f7d37-64b8-4d71-ac37-90849517e8c3 · outbound

This paper cites The Alignment Problem from a Deep Learning Perspective.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Alignment Problem from a Deep Learning Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.786686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.786686Z digest=sha256:9fc4befb74673d7b98647b327b52a3cb0a3a17a647658e4d29857f01f59b32e0

Observation 2dd13208-9c3f-41db-a050-f9db41382677 · outbound

This paper cites Training language models to follow instructions with human feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Training language models to follow instructions with human feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.791545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.791545Z digest=sha256:a95a4ba5579126209143af1df82fd1b7d6a271ff345f046130a982e3be085a16

Observation e8f00cd7-f125-4cc5-a484-0388ae6fc566 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.796035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.796035Z digest=sha256:52387090890a8d818fa32100f69bac3a718a7db4ccac222226d735bce0e0338c

Observation 484ec70b-f921-4e9e-983b-99a690623927 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Discovering Language Model Behaviors with Model-Written Evaluations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.800654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.800654Z digest=sha256:6a92fb45c5dafb7e91d84fade69308e67240a07ba50a68e65243700618e0e7dc

Observation e1cec0d2-b5a3-4ae4-8ed2-1b58c08b1be4 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.805283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.805283Z digest=sha256:fe1ac91bf160ce06eb1e425d8bda55536d6c63d3edac45945b357c380fdd95b0

Observation 6124cd75-fc2f-4c8a-8dd5-1d484253f19c · outbound

This paper cites On the Spectral Bias of Neural Networks.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes On the Spectral Bias of Neural Networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.809491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.809491Z digest=sha256:2d79b0c2bb826f8e65a4c465e1913246dfaf87ea7a9d40b70c0e8c5c52b2d5aa

Observation 67ff7e07-67ae-45d8-8463-9eea83e2a93e · outbound

This paper cites an unresolved cited work.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:18:01.163686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.814145Z digest=sha256:a44ed2c445aa5c317ff48bf1fadc01b708dc8339b76594297118d87a39741e66

Observation 0f81ba7d-5427-4d0d-ab0d-64da85f6dcc6 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.822017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.822017Z digest=sha256:29249b450be9a702ae6aa3e910270ac5028df34b0e9172b66ebf08ed89b38873

Observation 70b8a9c3-467a-4eff-a4ce-6e9bf5c5cdce · outbound

This paper cites Defining and Characterizing Reward Hacking.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Defining and Characterizing Reward Hacking

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.828256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.828256Z digest=sha256:db67d4b51bb4fae414d8c2a09cd185725d4d1fbaa3ca8dc74bd3347d0902f87a

Observation 8897593c-21db-4e5e-82de-8e61f6e842a1 · outbound

This paper cites Learning to summarize from human feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Learning to summarize from human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.845655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.845655Z digest=sha256:0fda0890a4fd4d5a428a87650fc66965bf6217233bd9f69331f666c4ec673bcb

Observation ff6f8909-8f9b-422f-a18f-1469f7c91ff8 · outbound

This paper cites Inoculation prompting: Eliciting traits from LLMs during training can suppress them at test-time, 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Inoculation prompting: Eliciting traits from LLMs during training can suppress them at test-time, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.850544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.850544Z digest=sha256:c50c9a4d19781de9acfbc90d4ff7571902b20f37192a92700dda92fad0ada32b

Observation 5c435b4a-93a8-44f1-a40f-6e524e070d3c · outbound

This paper cites Deep learning generalizes because the parameter-function map is biased towards simple functions.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Deep learning generalizes because the parameter-function map is biased towards simple functions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.855765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.855765Z digest=sha256:84eb63d8ff36dabd8011d6bfe50f473e4be354e5f640c7de5ba92e3fd22a67d1

Observation 4f4f9820-c79c-477a-9f39-3968f834b1a3 · outbound

This paper cites Inoculation prompting: Instructing LLMs to misbehave at train-time improves test-time alignment, 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Inoculation prompting: Instructing LLMs to misbehave at train-time improves test-time alignment, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.862922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.862922Z digest=sha256:9ddfa32191ef75a52f72d5218d85dd3e35d31da8a71be8d62d1b3135003cf3eb

Pith citing papers

No inbound Pith citation observations are available.