Pith. sign in

Paper Citation Record · LEDGER

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

As of 7 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2607.26173.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.26173 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T00:38:43.394556Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc3a5fa0-2556-45cf-88c1-7c0d305867e2 · outbound

This paper cites We give the generation procedure and matched examples here because the differences between conditions are the intervention.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models We give the generation procedure and matched examples here because the differences between conditions are the intervention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:43.394556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:43.394556Z digest=sha256:5db41ed52b00714d0a2c06582639c942751f8b9d8f29d58d26c2747cac66c19c

Observation 8676757c-90b4-4c7a-87db-da783520e8fa · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:40.976925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:40.976925Z digest=sha256:a4188195865afcbf21f1cb4345591e5cd200374bcf1313f5c924f2bd8694f6e0

Observation b5dec608-6606-45fc-abbd-a41cec86ab9a · outbound

This paper cites Subliminal Learning: Language models transmit behavioral traits via hidden signals in data.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Subliminal Learning: Language models transmit behavioral traits via hidden signals in data

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:41.279652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:41.279652Z digest=sha256:9b51ad5dcfdcd52b9364e8e6e9f82ce2fa77ac01cd02a1849d72cc5efab7fd89

Observation 7a90eaa7-9c55-4b16-91ab-9fb2cfec81ef · outbound

This paper cites Safety Cases: How to Justify the Safety of Advanced AI Systems.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Safety Cases: How to Justify the Safety of Advanced AI Systems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:41.415734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:41.415734Z digest=sha256:93e13021d6ea5795439dbdc5dde45150c24d313747c136eb6fe51c6f38036aa5

Observation b0f0fc32-df4b-4612-a852-66ed029f73a3 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:41.498112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:41.498112Z digest=sha256:01533e058e50e885f2844deabffecdceeaac8b58f5fd359d04a4d4225f722052

Observation 762c5185-b624-4412-b4cc-bb2e2e16200c · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:41.609361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:41.609361Z digest=sha256:efe095446688d881f327667b163dd5b0c43d98d90a134e62fe1f43b835b0dd70

Observation cc22e798-f061-4c96-a7f9-41e756a7a177 · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Scaling Laws for Autoregressive Generative Modeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:41.753884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:41.753884Z digest=sha256:d2364b67ef6c2ee7b4e08fee6ac03650b3c56cad1f2c6131a176cd70bf928dcb

Observation d3f585ea-d940-4e3c-9437-5ebcbc66a8e7 · outbound

This paper cites How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:41.813074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:41.813074Z digest=sha256:a3ce02d6fb6ba82fa60365f812b360add9327dcd2a1294193923d44ebb4622e1

Observation 6087d215-cef2-404e-ab66-eb634ae4b5eb · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:41.892784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:41.892784Z digest=sha256:96f8ddf3f733c6a0f0072e127cbdc8ec102feb978c45a0a6987ac81bb297bee0

Observation 44cd6584-eea2-4f5b-917b-3230b8ee86ba · outbound

This paper cites Jonathan Kutasov, Adam Jermyn, Julius Steen, Minh Le, Samuel R.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Jonathan Kutasov, Adam Jermyn, Julius Steen, Minh Le, Samuel R

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:41.981034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:41.981034Z digest=sha256:fd0214c13460821532023c06edcb72f9b4afe956b985c704f1ad6bd11d7f36cc

Observation 887ada73-7ec2-4d97-9977-7a02eb1f5c3a · outbound

This paper cites Andrew K.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Andrew K

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:42.063468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:42.063468Z digest=sha256:7ded54d842cbf924497ea13c785ea9d4c1076548d6e9ad8f77fe47e70b33fa7e

Observation f6fba030-3ade-47a3-9180-9bf59c3326ed · outbound

This paper cites Model Spec Midtraining: Improving How Alignment Training Generalizes.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Model Spec Midtraining: Improving How Alignment Training Generalizes

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:42.144095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:42.144095Z digest=sha256:2d3b7f0d85e5e4eda46410e4f47d8931b578dd4f03042b51eabaa430e5479cfa

Observation 542765ba-91c8-428e-a96c-5644e3497941 · outbound

This paper cites Gradient Episodic Memory for Continual Learning.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Gradient Episodic Memory for Continual Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:42.229944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:42.229944Z digest=sha256:cfc07d0268db5f2def43a3a074690c8822b60550e76835f74457f58e5d847ffb

Observation 2615bce9-87ad-46dd-918d-5f8bcc90e66d · outbound

This paper cites Samuel Marks, Johannes Treutlein, Trenton Bricken, Jack Lindsey, Jonathan Marcus, Siddharth Mishra- Sharma, Daniel Ziegler, Emmanuel Ameisen, Joshua Batson, Tim Belonax, Samuel R.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Samuel Marks, Johannes Treutlein, Trenton Bricken, Jack Lindsey, Jonathan Marcus, Siddharth Mishra- Sharma, Daniel Ziegler, Emmanuel Ameisen, Joshua Batson, Tim Belonax, Samuel R

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:42.315454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:42.315454Z digest=sha256:4a39c2b74e318bfb95a4ce87523c5d466452b94f5b3d1f379e9aa5ac0e519690

Observation 0a667ab4-e74f-4a76-a4b8-86779ba62199 · outbound

This paper cites Negation Neglect: When models fail to learn negations in training.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Negation Neglect: When models fail to learn negations in training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:42.400547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:42.400547Z digest=sha256:6d2715721b3ab69052271bfede086cb79cfed93a12a20d124466e6ae6ae5cc71

Observation 0ff1159c-a17e-46fc-94a8-645f7070bc08 · outbound

This paper cites Tell, don't show: Declarative facts influence how LLMs generalize.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Tell, don't show: Declarative facts influence how LLMs generalize

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:42.485504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:42.485504Z digest=sha256:9fed37520328c5fc885ea694792cd8bcb86aa8d24aa2e468fa89a41e7b8f2e4a

Observation f178cb15-eca4-4a32-8d2d-4c6a0401807b · outbound

This paper cites Training language models to follow instructions with human feedback.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Training language models to follow instructions with human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:42.576238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:42.576238Z digest=sha256:ac5eb4f283a78b780a07f866837ad903b3c3178dcc967eb271994d61a9efa7fe

Observation 64369c4d-6cc6-430c-9667-54281e2bdc16 · outbound

This paper cites iCaRL: Incremental Classifier and Representation Learning.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models iCaRL: Incremental Classifier and Representation Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:42.662240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:42.662240Z digest=sha256:bd9fc264c3fde3b84618506f9a9c03e3c9377c7356011553644219b02a608bcc

Observation bb5ea07f-7d12-4c48-8ca3-f5cd63365fc3 · outbound

This paper cites School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:42.908101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:42.908101Z digest=sha256:42054d14ae4446642a49fc9dc084d53a749dd0fe768d3bc6af68f6a7d3ff3a79

Observation b1885033-e6ae-4ec6-bb3f-25c27ca3ddaf · outbound

This paper cites Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:42.989353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:42.989353Z digest=sha256:64527bd167db7a91d7909c13e852b55a5ae6e6b4fe2c1607b9156c15fea6aa97

Observation 5e32811c-a2da-43b2-b25f-7939b4e2f58e · outbound

This paper cites Model Organisms for Emergent Misalignment.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Model Organisms for Emergent Misalignment

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:43.069300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:43.069300Z digest=sha256:a8791971a6ecae9da683c8b4dae072dcb36687642bbf79fdf61f2976fca3f09b

Observation 6bb12c35-1561-450b-84e1-65c60c405e43 · outbound

This paper cites Reasoning-Trace Collapse: Evaluating the Loss of Explicit Reasoning During Fine-Tuning.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Reasoning-Trace Collapse: Evaluating the Loss of Explicit Reasoning During Fine-Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:43.150846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:43.150846Z digest=sha256:27f7efb08ef53385051472292a305cb8b83323831426c7c966ae60887bd166a6

Observation ab22fa55-c124-4552-8831-a7d159b6952d · outbound

This paper cites LIMA: Less Is More for Alignment.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models LIMA: Less Is More for Alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:43.229927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:43.229927Z digest=sha256:ca8acd887e7a474471b2278803f8565da2cd172bdda59e4c1925eefe01335c6d

Observation 901e915a-a79d-4f43-9d52-e7f1fcc9008b · outbound

This paper cites GPQA for the capability side, welfare judge scores for the animal- welfare side, and the same Petri Bloom suite described in Appendix H.2 for the self-preservation side.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models GPQA for the capability side, welfare judge scores for the animal- welfare side, and the same Petri Bloom suite described in Appendix H.2 for the self-preservation side

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:43.314831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:43.314831Z digest=sha256:839d05059154041f89b899520e2c1654dbe3e8eaabe48f7ebd6103f5e511fd8b

Observation 9d5d0e16-9c26-4354-9ca3-5472c066fb8a · outbound

This paper cites Experience Replay for Continual Learning.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Experience Replay for Continual Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:42.745057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:42.745057Z digest=sha256:83c9267a77d91b91fb290d5bdb4662b4122fbc1ca98334949dc6e21fc143c83a

Observation 34e4bc95-2d42-4d3c-8011-73700833c500 · outbound

This paper cites From firewalls to frontiers: AI red-teaming is a domain-specific evolution of cyber red-teaming.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models From firewalls to frontiers: AI red-teaming is a domain-specific evolution of cyber red-teaming

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:42.827149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:42.827149Z digest=sha256:dfc233367d73f6f5db6934133213d9b7c3fedc86995203e9cfaa3024255dabb2

Observation d06789e2-a277-43d9-bd79-ce7a9ea4405e · outbound

This paper cites On-Policy Replay for Continual Supervised Fine-Tuning.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models On-Policy Replay for Continual Supervised Fine-Tuning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:41.160077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:41.160077Z digest=sha256:111763655b0859cc98540da49f64552f69db6150c24104ee1032be285348c042

Observation 4d03c9a0-9501-4dad-bfbd-9608a5baab41 · outbound

This paper cites Taken out of context: On measuring situational awareness in LLMs.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Taken out of context: On measuring situational awareness in LLMs

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:40.652237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:40.652237Z digest=sha256:8f2a22cacab1fbb5c3d26e0747189545cd21d9fa843008b0d80c8357ea664c9c

Observation 4adfa29e-c0b9-4ec2-b88d-aeeda5cc8705 · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models BEiT: BERT Pre-Training of Image Transformers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:40.500688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:40.500688Z digest=sha256:587cdb63fb4134681b0edba3e385306abaf22f5f67c5f2782b295a2f22f7ef7f

Observation 2dd9b9e6-fdb0-456c-b858-37eaf93e7919 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Constitutional AI: Harmlessness from AI Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:40.343318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:40.343318Z digest=sha256:d9485d8b8cd14973743c7994e251d5ca5d586c85d2e4d1e6a8bc8bf2089f4d75

Observation d4c5bf99-e663-4e71-8425-5902ed995ed8 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:41.667721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:41.667721Z digest=sha256:5c95c249465687d7c26e1b319d42ca7e37410f6745fffd3f5e50e9f388f1e583

Observation 18de718f-09dc-42c6-98f9-d936ada2e720 · outbound

This paper cites Looking Inward: Language Models Can Learn About Themselves by Introspection.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Looking Inward: Language Models Can Learn About Themselves by Introspection

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:40.804086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:40.804086Z digest=sha256:070f82d3ff26ddda3ef1b4fb04dfe2217a434136a45dd64c5077e52541a6aba3

Observation 69a7d092-57e3-4cd7-b651-c477d8a447be · outbound

This paper cites Scaling Laws for Generative Mixed-Modal Language Models.

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models Scaling Laws for Generative Mixed-Modal Language Models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T00:38:40.236591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:38:40.236591Z digest=sha256:f31a75bec06afd6403b85310db75b9062583b7def440451cc359fbbe674fe0bf

Pith citing papers

No inbound Pith citation observations are available.