Pith. sign in

Paper Citation Record · LEDGER

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes

As of 14 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.10209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10209 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:17:59.862922Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a940407a-d1b7-4435-90f9-b189062a0605 · outbound

This paper cites Concrete Problems in AI Safety.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Concrete Problems in AI Safety

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.686636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.686636Z digest=sha256:ec24d9d589a4aa2b97f6d9b7cc7b29818d611e1a3b35006154e784a153795d76

Observation 1bc19727-d728-4384-be6e-5f8a6052a8a3 · outbound

This paper cites Measuring political bias in Claude.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Measuring political bias in Claude

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:01.210935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.693517Z digest=sha256:56154e3b9c9af7c5b3614086ae0465166102c225049bdb466afa0e3bf284c675

Observation 3ab7e301-3aa1-4910-8a42-51fbe73230b0 · outbound

This paper cites The Internal State of an LLM Knows When It's Lying.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Internal State of an LLM Knows When It's Lying

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.698106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.698106Z digest=sha256:7cc0bb44d2829721fa97e76ccd02a7737d1c1d0a769235aea2b9d3bf89a62613

Observation 0728dbc1-5e0a-4ce9-82c5-b695b11f014e · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.702774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.702774Z digest=sha256:e4691ffb2c14c1e6aaba2a496997ae803eee85dd111b517f6b0af624a09dcfd3

Observation a5248c68-d469-4d80-9043-a6a7c99f8755 · outbound

This paper cites Boerner, Stephen Deems, Thomas R.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Boerner, Stephen Deems, Thomas R

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.708085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.708085Z digest=sha256:ac1cc266076c6caa5288b43eec5b2994a3d9b8ccf9e9bd844d41eefd9154f4a2

Observation ad5c8121-5601-44e6-8bce-716bff1b6c15 · outbound

This paper cites Discovering Latent Knowledge in Language Models Without Supervision.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Discovering Latent Knowledge in Language Models Without Supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.712832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.712832Z digest=sha256:a256f9283534cd1998d16e178a5cfdd7ab3538fa4a23d1b06bc5e510d9ae2be6

Observation d314d827-13af-455e-ad6c-9ebf322be669 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.721800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.721800Z digest=sha256:c66915c4cc27c1ee6bcfcba04d85e1ed62e06e1262f958c41e3420ec8f9e479f

Observation 1532d86c-f36f-4370-bcec-d38badff26c4 · outbound

This paper cites Eliciting latent knowledge: How to tell if your eyes deceive you.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Eliciting latent knowledge: How to tell if your eyes deceive you

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:18:01.186919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.729878Z digest=sha256:27395cd3dc611b8b616ca4860e12d363a7aee10593d9e136de1ea312312c60c3

Observation ad8e74f6-8a6a-4dca-aa03-5f6944d8fec8 · outbound

This paper cites Deep reinforcement learning from human preferences.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Deep reinforcement learning from human preferences

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.734216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.734216Z digest=sha256:c87098ab447eab935f0eb4b220eeac420c6ba6d51a68854615c53fea06063aa2

Observation c4face5c-2283-4977-9f4f-19466581e0e7 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.739636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.739636Z digest=sha256:617f5aabe20e0ab9fe72873af6eb983469b738a0459f03669322b8b72097b19e

Observation 8deeab91-14d7-4f38-a260-c36fdde05c1d · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes QLoRA: Efficient Finetuning of Quantized LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.746751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.746751Z digest=sha256:5ff1a5ec7bb97a9f05c3097f05bbe8aa0d885a79a50e8c6ae7c28b8a7997089a

Observation 1660b117-16f9-414a-a927-d7ebe0524aea · outbound

This paper cites On the Relationship between Truth and Political Bias in Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes On the Relationship between Truth and Political Bias in Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.751481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.751481Z digest=sha256:bf37d5443559b3d22094971e4148de56afde48eb95b6b286cef6ee704a24ff73

Observation a3ef1fa3-574b-4cfa-90a8-f5f173373f36 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes LoRA: Low-Rank Adaptation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.756004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.756004Z digest=sha256:08f00a884f368a22bb013b56bf1f367be4e573c1fa0ef9411a7a4591acccb38f

Observation 4ccd811d-2e17-4c4d-a8a6-218685ff43cb · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.760327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.760327Z digest=sha256:f3b1c385dfe65ea0efdd0f144b9624a118e702ec890eacbb0dca182ed0313a40

Observation 07504f46-8113-4612-b4ea-b006646da988 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Language Models (Mostly) Know What They Know

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.767334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.767334Z digest=sha256:fdb587ee2d5d100f96d4cd30166212bd5d2971b1256799461ac67a3be6fc8b04

Observation 2ec66086-0e3d-43db-936c-d436265fe252 · outbound

This paper cites SGD on Neural Networks Learns Functions of Increasing Complexity.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes SGD on Neural Networks Learns Functions of Increasing Complexity

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.772617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.772617Z digest=sha256:fa0eeeddb024e54b522d829c7fc1cea42de4c977d6b11b637ca104c9261673c8

Observation 2e957e4c-300d-4875-9541-0ca9b905fcd1 · outbound

This paper cites Natural emergent misalignment from reward hacking in production RL , 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Natural emergent misalignment from reward hacking in production RL , 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.777339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.777339Z digest=sha256:38750e19736c7cea667066f5d524303bcea0fc1ace593a767b1a10320af7fbf7

Observation 768a663e-8244-4445-a707-0a90224fa3fe · outbound

This paper cites Categorizing Variants of Goodhart's Law.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Categorizing Variants of Goodhart's Law

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.782461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.782461Z digest=sha256:646b256982473577ae6d364c1c25024c9374123841d754b2456e94fc2195b84e

Observation 745f7d37-64b8-4d71-ac37-90849517e8c3 · outbound

This paper cites The Alignment Problem from a Deep Learning Perspective.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Alignment Problem from a Deep Learning Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.786686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.786686Z digest=sha256:2905e36eaf8c09acf23c958ad23262179bccd104da80b4894bb6de3f60aef706

Observation 2dd13208-9c3f-41db-a050-f9db41382677 · outbound

This paper cites Training language models to follow instructions with human feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Training language models to follow instructions with human feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.791545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.791545Z digest=sha256:ccf52594ae5961b57b83e7d694eca25d585ad553d20d5302e0c730ef7a7de283

Observation e8f00cd7-f125-4cc5-a484-0388ae6fc566 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.796035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.796035Z digest=sha256:7bae2da618e828acf79b26638b13523d869d08864320f8062982216aa0e3648b

Observation 484ec70b-f921-4e9e-983b-99a690623927 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Discovering Language Model Behaviors with Model-Written Evaluations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.800654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.800654Z digest=sha256:6f1c1ee4d184b3dc033073d057e70f85b878ab8d46358a51b7a1a64571133488

Observation e1cec0d2-b5a3-4ae4-8ed2-1b58c08b1be4 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.805283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.805283Z digest=sha256:74dd9531afad0af8dfcd266d31607ae77d22151bba9c9627d7a7e61f31c0f583

Observation 6124cd75-fc2f-4c8a-8dd5-1d484253f19c · outbound

This paper cites On the Spectral Bias of Neural Networks.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes On the Spectral Bias of Neural Networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.809491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.809491Z digest=sha256:611e2efed5e1f170700aea34a76527b0722ef5a33ea7d1d908f56a0558922642

Observation 67ff7e07-67ae-45d8-8463-9eea83e2a93e · outbound

This paper cites an unresolved cited work.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:18:01.163686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-14T04:17:59.814145Z digest=sha256:6e349a6076a93d08135ee8b36ad5be2c0d4a7dc882e21ac9512ade3ed0c8aa7a

Observation 0f81ba7d-5427-4d0d-ab0d-64da85f6dcc6 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.822017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.822017Z digest=sha256:13c2c486a1d4d8ad0806236267da43cb57778b47efb0c7c8cba93aa92288ea01

Observation 70b8a9c3-467a-4eff-a4ce-6e9bf5c5cdce · outbound

This paper cites Defining and Characterizing Reward Hacking.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Defining and Characterizing Reward Hacking

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.828256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.828256Z digest=sha256:f8dee845150a483248966f119bb5cc38ede8fcf6ebbede47fd346495290fbfe8

Observation 8897593c-21db-4e5e-82de-8e61f6e842a1 · outbound

This paper cites Learning to summarize from human feedback.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Learning to summarize from human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.845655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.845655Z digest=sha256:4aff1e74d41ee1a9893baf9f842de3279835e39c676ff9e934fac2490951028e

Observation ff6f8909-8f9b-422f-a18f-1469f7c91ff8 · outbound

This paper cites Inoculation prompting: Eliciting traits from LLMs during training can suppress them at test-time, 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Inoculation prompting: Eliciting traits from LLMs during training can suppress them at test-time, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.850544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.850544Z digest=sha256:9190bcedf02100fa90ac1d8b8b6a64e60fbbbfc19b79da9b9c114cb6fe676b44

Observation 5c435b4a-93a8-44f1-a40f-6e524e070d3c · outbound

This paper cites Deep learning generalizes because the parameter-function map is biased towards simple functions.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Deep learning generalizes because the parameter-function map is biased towards simple functions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.855765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.855765Z digest=sha256:1374fb6768c6de37343f5ee4ef44bfed8d25dce7a58ac10415db594057aabfce

Observation 4f4f9820-c79c-477a-9f39-3968f834b1a3 · outbound

This paper cites Inoculation prompting: Instructing LLMs to misbehave at train-time improves test-time alignment, 2025.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes Inoculation prompting: Instructing LLMs to misbehave at train-time improves test-time alignment, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.862922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.862922Z digest=sha256:8e06d4d81547699f19466757c5a9d6b64b9d2cac279b36fda2c7855b11040708

Pith citing papers

No inbound Pith citation observations are available.