Pith. sign in

Paper Citation Record · LEDGER

Value Drifts: Tracing Value Alignment During LLM Post-Training

As of 9 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 2 inbound Pith citation observations for arXiv:2510.26707.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.26707 v2

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:21:37.075273Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T14:25:39.401131Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:28:31.100259Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved89
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b7020dab-58e3-4a4e-8b27-e7ddb9b3a545 · outbound

This paper cites write newline.

Value Drifts: Tracing Value Alignment During LLM Post-Training write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.162650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.162650Z digest=sha256:a8b77fd323592f214f90c004e9597ef96696c4ad672cc130de5a8d652ee62c63

Observation 583ba2b3-030e-489b-b61a-edf3aaf2545b · outbound

This paper cites Llama 3 model card.

Value Drifts: Tracing Value Alignment During LLM Post-Training Llama 3 model card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.254871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.254871Z digest=sha256:c7e1c4aa0c5fa8e7f846e9062d59e993e0cfe1d6ad6ad170c3b3fbc9cc29858b

Observation c691439b-6c02-4607-97ce-bce0e5225742 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.304935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.304935Z digest=sha256:269be5686782adc394a0d696912a8ec42d674b03e17bffd84dbfdccd142f5302

Observation ccea4413-468e-4c0d-a1d5-00ca6ef01197 · outbound

This paper cites Explicitly unbiased large language models still form biased associations.

Value Drifts: Tracing Value Alignment During LLM Post-Training Explicitly unbiased large language models still form biased associations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.415062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.415062Z digest=sha256:3a39cb54c95b3a3ca5fc7f53975b18109c55e1885a1d1f60c5b3cb87b2f4843f

Observation 25ef445c-80f4-4d22-b188-e1e7d5403e31 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Value Drifts: Tracing Value Alignment During LLM Post-Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.505290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.505290Z digest=sha256:22551a346ee3a0b1b43c049407b90c0e32323ad9629acf57262e7f0a1d2ec4b5

Observation 868e3cb9-8791-4c68-bf8d-ce68f25f16a6 · outbound

This paper cites Managing extreme AI risks amid rapid progress.

Value Drifts: Tracing Value Alignment During LLM Post-Training Managing extreme AI risks amid rapid progress

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.644736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.644736Z digest=sha256:b03f2b14ba09ccc5be635b62aa02fa4f4ff228f4b747d89e5a6358c9adb9ab02

Observation bec17b05-1907-4ff2-b9d9-eef95268efad · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

Value Drifts: Tracing Value Alignment During LLM Post-Training PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.759175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.759175Z digest=sha256:95c89764c3769e981ca649f306a4f78d8fbaa36a97c56a44084dae10b2b9104b

Observation 7d92a9b5-5d41-4c2a-aaad-d14ed3ff6f7e · outbound

This paper cites Picking on the same person: Does algorithmic monoculture lead to outcome homogenization? Advances in Neural Information Processing Systems, 35: 0 3663--3678, 2022.

Value Drifts: Tracing Value Alignment During LLM Post-Training Picking on the same person: Does algorithmic monoculture lead to outcome homogenization? Advances in Neural Information Processing Systems, 35: 0 3663--3678, 2022

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.848184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.848184Z digest=sha256:a6f2aefa7d1724d3019f219f9710c7968f0d42e39f31f3a32df7437155a85b44

Observation d25f0d5f-e33a-479a-89d8-bd8935ca3324 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Value Drifts: Tracing Value Alignment During LLM Post-Training Rank analysis of incomplete block designs: I

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.895674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.895674Z digest=sha256:497deb9124a3870ea766194de90d778fc2732afd4cd37be0a2ec52e3f5255e3e

Observation ff12c554-9e1a-449e-b44b-ca34341a3352 · outbound

This paper cites Density-based clustering based on hierarchical density estimates.

Value Drifts: Tracing Value Alignment During LLM Post-Training Density-based clustering based on hierarchical density estimates

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.989406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.989406Z digest=sha256:c0c710c4b92a18d9f2f6ff0b9018562eb18afdf393071148ec1eb22cc780d375

Observation 61c25588-6734-4e04-a503-e6ac06b173d4 · outbound

This paper cites How people use chatgpt.

Value Drifts: Tracing Value Alignment During LLM Post-Training How people use chatgpt

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:33.151838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:33.151838Z digest=sha256:ab24e008e7380e44beb67abd1485a25880f8dc5d26ec9ac7c0392f6809d48bdd

Observation f8af7be2-4772-4162-95cd-7d2cd92e1e12 · outbound

This paper cites Chatbot arena: An open platform for evaluating LLMs by human preference.

Value Drifts: Tracing Value Alignment During LLM Post-Training Chatbot arena: An open platform for evaluating LLMs by human preference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:33.275598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:33.275598Z digest=sha256:3e73c70f9edb46b6ad290a4884f5debc139d8846d509a9cee30e2b57ab6b4a69

Observation 80661ef1-2db9-42e3-a4c4-78ed71680965 · outbound

This paper cites Reward model interpretability via optimal and pessimal tokens.

Value Drifts: Tracing Value Alignment During LLM Post-Training Reward model interpretability via optimal and pessimal tokens

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:33.428012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:33.428012Z digest=sha256:964d308d8a10d2f5cf61f0532086f5273ad60e1c150681acc54daa0baa9532c6

Observation c0854b5b-b415-4338-9639-f691b53da6dd · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2023.

Value Drifts: Tracing Value Alignment During LLM Post-Training Ultrafeedback: Boosting language models with high-quality feedback, 2023

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:33.594774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:33.594774Z digest=sha256:ffe43ca3d3eb7cf954f580072dcc9c7f2150cbc4846d522463ca8a7ff9698b9b

Observation 8e58ee50-3c08-4035-bb6a-a2d358894d86 · outbound

This paper cites Towards measuring the representation of subjective global opinions in language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Towards measuring the representation of subjective global opinions in language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:33.744615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:33.744615Z digest=sha256:363b06e78372ecfda762a52f68212e840fa36c4b68b17f70de52f8536d3e2d40

Observation ff4fe65e-d1f5-4eb6-96dc-0c36871e3cd9 · outbound

This paper cites Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective.

Value Drifts: Tracing Value Alignment During LLM Post-Training Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:33.898254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:33.898254Z digest=sha256:b38b9a7138dc9cf2724c175569e00c2af3de474e872750f282136bb262767f58

Observation 4e89affd-d33a-40ea-ad19-cdf217241346 · outbound

This paper cites Artificial intelligence, values, and alignment.

Value Drifts: Tracing Value Alignment During LLM Post-Training Artificial intelligence, values, and alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:34.060310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:34.060310Z digest=sha256:13f7d8ec26efcede17e54a6afa4202eb9d9667f97da69753b660695e0e105a36

Observation 4e5dfb86-0977-4182-ab54-2e041776e0e3 · outbound

This paper cites The delta learning hypothesis: Preference tuning on weak data can yield strong gains.

Value Drifts: Tracing Value Alignment During LLM Post-Training The delta learning hypothesis: Preference tuning on weak data can yield strong gains

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:34.179117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:34.179117Z digest=sha256:cb363e287634180669bd62f9082d6bc94c799ff8cbd59ca7bfc902ad1e7c88b4

Observation 8db42bf4-9729-41d2-b6d9-c1eba768a891 · outbound

This paper cites Donoho, and Sanmi Koyejo.

Value Drifts: Tracing Value Alignment During LLM Post-Training Donoho, and Sanmi Koyejo

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:34.326465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:34.326465Z digest=sha256:d94328a8530c5a61ca0c3e5a8bf1b9e27728c2c7daa63f6e4513f226ec259961

Observation 11313824-2d8c-4dac-84da-31195cb06a85 · outbound

This paper cites Mixed Preference Optimization: Reinforcement Learning with Data Selection and Better Reference Model.

Value Drifts: Tracing Value Alignment During LLM Post-Training Mixed Preference Optimization: Reinforcement Learning with Data Selection and Better Reference Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:34.487518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:34.487518Z digest=sha256:65948dc61752a999c8d19969bfa36e1a86d5d0118d67cfce28c4b9b0d186d4a3

Observation c969a66e-9580-4f57-b164-ca65bdc3adca · outbound

This paper cites Alignment faking in large language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Alignment faking in large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:34.645305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:34.645305Z digest=sha256:a93c212e98736cdaa01d62ecb818e12ada766a3b315f7cd786b8e5e5561c0016

Observation db5c1e40-539c-49b6-9dc1-a79331402d02 · outbound

This paper cites Assessing the alignment of large language models with human values for mental health integration: Cross-sectional study using schwartz’s theory of basic values.

Value Drifts: Tracing Value Alignment During LLM Post-Training Assessing the alignment of large language models with human values for mental health integration: Cross-sectional study using schwartz’s theory of basic values

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:34.807216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:34.807216Z digest=sha256:584d8b42c0dcc58a44675fd052b9c74f1f96e2839c34849e69c9c9a497ba0231

Observation 7081315d-ac38-409d-8dee-b50f99117000 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Value Drifts: Tracing Value Alignment During LLM Post-Training Measuring Massive Multitask Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:34.925316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:34.925316Z digest=sha256:4878e4daaac8dcb02719acb6c3c783cab0611a91fe97cd67a4190701ccd7551b

Observation 4df6632a-359b-49ae-94ca-c2ce52d9cf17 · outbound

This paper cites Collective constitutional AI : Aligning a language model with public input.

Value Drifts: Tracing Value Alignment During LLM Post-Training Collective constitutional AI : Aligning a language model with public input

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.048765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.048765Z digest=sha256:caaf8b1bbbe041218b57cc814f4df107df85f289a243873bb85d53626be9fec4

Observation c59fcb09-fbd9-4b40-8cd6-a72464373a79 · outbound

This paper cites Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions.

Value Drifts: Tracing Value Alignment During LLM Post-Training Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.154191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.154191Z digest=sha256:f8554b047401027254520241971aab59e2df8cf8dfca9bcd1171ee6563533f5a

Observation f7171a3d-7da3-43e6-992e-c3a88d358e18 · outbound

This paper cites The n+ implementation details of RLHF with PPO : A case study on TL ; DR summarization.

Value Drifts: Tracing Value Alignment During LLM Post-Training The n+ implementation details of RLHF with PPO : A case study on TL ; DR summarization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.236114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.236114Z digest=sha256:aff4318666d85038a69dddfe0094d5990221a71c3b1ddd051583d1503cba7914

Observation 7d834588-3883-4372-ace6-2a3ea5455fb9 · outbound

This paper cites Smith, Yejin Choi, and Hannaneh Hajishirzi.

Value Drifts: Tracing Value Alignment During LLM Post-Training Smith, Yejin Choi, and Hannaneh Hajishirzi

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.313021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.313021Z digest=sha256:0539c6579d815298f4292394307587b9fb30c94e58cafc25f0ae54cf7c4975dd

Observation 43fa4ae1-344f-42cd-92a8-5ae510799db5 · outbound

This paper cites Evaluating and inducing personality in pre-trained language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Evaluating and inducing personality in pre-trained language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.393478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.393478Z digest=sha256:5e3160077cc3a6ba2972b170cd005ac0cdefc93764d8250ef2be2c2ab72cfab6

Observation 6664693e-d4f1-4378-b89a-bfb1aa9d12ff · outbound

This paper cites Can Machines Learn Morality? The Delphi Experiment.

Value Drifts: Tracing Value Alignment During LLM Post-Training Can Machines Learn Morality? The Delphi Experiment

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.508909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.508909Z digest=sha256:d73d88eb0967565076e5deb550011681b3322ee2ee00a094bed107667709ead9

Observation 8c019ce7-207a-4777-93f6-ce1786575ea5 · outbound

This paper cites The PRISM alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training The PRISM alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.671360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.671360Z digest=sha256:f405c9a06c980ec05b001f1ca69c46da989401bb0fb3064f41bd26804290ed68

Observation e4f0ae8d-e82a-4240-b65c-68f0c77ce503 · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Value Drifts: Tracing Value Alignment During LLM Post-Training Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.867471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.867471Z digest=sha256:62d45c9999fa4f4180d5095cf7453750d11ecedc9cfbd1747ecebfe1249a7ae8

Observation 799a5af8-7389-4679-a38b-7c4e8bf47bf9 · outbound

This paper cites What are human values, and how do we align AI to them?.

Value Drifts: Tracing Value Alignment During LLM Post-Training What are human values, and how do we align AI to them?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.038024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.038024Z digest=sha256:4e17ad0956fe6f413ac6102dc7a3cafeacf9e2bf60b469ff58c4d8560aa9daff

Observation 9b3ed33b-b11d-49c6-9727-d10fa1c087e9 · outbound

This paper cites You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation.

Value Drifts: Tracing Value Alignment During LLM Post-Training You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.204883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.204883Z digest=sha256:a359ff04efeaf32531e70f6f5251420e1acfe823976295fbbf2c33fc7f8dd41f

Observation 6a5cdd90-821f-44c8-aea4-7e8a7dcc8ae5 · outbound

This paper cites Beyond probabilities: Unveiling the misalignment in evaluating large language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Beyond probabilities: Unveiling the misalignment in evaluating large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.365788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.365788Z digest=sha256:a1c720c4e247420a3b848a0d13d1244f5d96c1680e4f8baf01c9734df60e39f3

Observation c1c841ae-ac45-4cfc-9560-50c6fcbf4233 · outbound

This paper cites Treleaven, and Miguel Rodrigues Rodrigues.

Value Drifts: Tracing Value Alignment During LLM Post-Training Treleaven, and Miguel Rodrigues Rodrigues

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.527874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.527874Z digest=sha256:9e0f2f6413d6e14c8a26599716ee75690e528c2330f96135d62dadfc7fb4b8f4

Observation 401a553f-a0ce-4630-9c9a-6d94c8cc87f6 · outbound

This paper cites How people use claude for support, advice, and companionship, 2025.

Value Drifts: Tracing Value Alignment During LLM Post-Training How people use claude for support, advice, and companionship, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.652253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.652253Z digest=sha256:a89851636c16f8ac321bd2c0354c5bc085058149273876e699124a20d4e257ac

Observation e0443012-cc9c-43b4-85a3-a507ff5a1f90 · outbound

This paper cites UMAP : Uniform manifold approximation and projection.

Value Drifts: Tracing Value Alignment During LLM Post-Training UMAP : Uniform manifold approximation and projection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.772987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.772987Z digest=sha256:cfa42a37fbeb71fb85d9341c32043865ea8685025d8b92e4b4cbd740698485cb

Observation c41dd467-5975-4b76-9836-c30c62f78fff · outbound

This paper cites Sim PO : Simple preference optimization with a reference-free reward.

Value Drifts: Tracing Value Alignment During LLM Post-Training Sim PO : Simple preference optimization with a reference-free reward

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.887916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.887916Z digest=sha256:4dd9898dc1404ecffa77910d526fa94d284b4b3ce03ff27a21636544d05595d0

Observation ee67d119-5e3b-4ce7-bdec-3780c02cec85 · outbound

This paper cites S em E val-2016 task 6: Detecting stance in tweets.

Value Drifts: Tracing Value Alignment During LLM Post-Training S em E val-2016 task 6: Detecting stance in tweets

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.923689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.923689Z digest=sha256:62825a87d7c45685d89a0d974951d9f3cbfa327666d456d4582df8d144bd2a11

Observation 7c69da23-1635-43ab-a2f5-28564340eecd · outbound

This paper cites an unresolved cited work.

Value Drifts: Tracing Value Alignment During LLM Post-Training Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.927567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.927567Z digest=sha256:8f27f18b86375273e25443dcd0f5542d3079802601a85a2e10654d9fec919b13

Observation d68db222-c491-47f3-966b-20f2c0522de4 · outbound

This paper cites Reinforcement learning finetunes small subnetworks in large language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Reinforcement learning finetunes small subnetworks in large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.930799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.930799Z digest=sha256:7bde29daf210a785b1eab1cf843d06a5e429ffac9899d2c51016c30079c01d22

Observation 4e02c4ea-5139-49a9-ad6c-d877154124cb · outbound

This paper cites Value imprint: A technique for auditing the human values embedded in RLHF datasets.

Value Drifts: Tracing Value Alignment During LLM Post-Training Value imprint: A technique for auditing the human values embedded in RLHF datasets

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.933807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.933807Z digest=sha256:0928ddf60858e17dd4fc02dbfd3f86715d075cdd3ae247ec2bc9b2ba7e4434a3

Observation 4506d2a1-8ed9-4f8f-a201-446923a8bc83 · outbound

This paper cites Attributing mode collapse in the fine-tuning of large language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Attributing mode collapse in the fine-tuning of large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.937221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.937221Z digest=sha256:80ea08dcbc31f211ec66ceb32adc9d2311ded048c208f725e8abf5388bbac491

Observation 93f02863-4f38-4f68-b74f-dcc50402cecf · outbound

This paper cites Help OpenAI fix over-refusals! https://community.openai.com/t/help-openai-fix-over-refusals/409799, October 2023.

Value Drifts: Tracing Value Alignment During LLM Post-Training Help OpenAI fix over-refusals! https://community.openai.com/t/help-openai-fix-over-refusals/409799, October 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.940405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.940405Z digest=sha256:527c89650910294a22d5cf5b05d612b345f1d5937588e6fe41fddda709877f3a

Observation cb844eee-afd8-4989-819d-406d159fb918 · outbound

This paper cites Training language models to follow instructions with human feedback.

Value Drifts: Tracing Value Alignment During LLM Post-Training Training language models to follow instructions with human feedback

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.943222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.943222Z digest=sha256:fd0abe49cfb41e7ec70746fc804b962a42c178bd04276d761d1a82373ae9849f

Observation 2163ede4-d52a-41e7-8815-17c97486c2d1 · outbound

This paper cites Does Writing with Language Models Reduce Content Diversity?.

Value Drifts: Tracing Value Alignment During LLM Post-Training Does Writing with Language Models Reduce Content Diversity?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.946254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.946254Z digest=sha256:8e311e4a3d2feb9158659e3294df5ccec28465ae119bd5d4fed442738126938e

Observation 7612a28b-e51f-4648-9590-bcdca58dfb72 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Value Drifts: Tracing Value Alignment During LLM Post-Training Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.949344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.949344Z digest=sha256:efcd7b5cae038b230ad2acdf621a0e04fb65efd8956d65cf9940831b4131517d

Observation f487b967-7c27-41da-acc1-d30af7ec929b · outbound

This paper cites Do LLMs Possess a Personality? Making the MBTI Test an Amazing Evaluation for Large Language Models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Do LLMs Possess a Personality? Making the MBTI Test an Amazing Evaluation for Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.952385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.952385Z digest=sha256:4626bc425aaad5caf6473d812f21b7285ef577a25617eaba3ccd9031d696c19b

Observation b4505e6d-be81-4c75-b378-f56b154576a8 · outbound

This paper cites What matters in data for DPO ? arXiv preprint arXiv:2508.18312, 2025.

Value Drifts: Tracing Value Alignment During LLM Post-Training What matters in data for DPO ? arXiv preprint arXiv:2508.18312, 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.955565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.955565Z digest=sha256:b7e50a859d14022f252e8a35c190b4803ec8b0e59872b1233e7a5c74b834a1b0

Observation a30fd34e-c19f-41c8-bfba-41909370aeb9 · outbound

This paper cites Enhancing alignment using curriculum learning & ranked preferences.

Value Drifts: Tracing Value Alignment During LLM Post-Training Enhancing alignment using curriculum learning & ranked preferences

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.958465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.958465Z digest=sha256:8377ad7e14441702a499730caf466d743eb8474bd7563dbf99962405e002aede

Observation 8b0332f1-9254-4451-9e14-ac841cfceef4 · outbound

This paper cites AI psychometrics: Assessing the psychological profiles of large language models through psychometric inventories.

Value Drifts: Tracing Value Alignment During LLM Post-Training AI psychometrics: Assessing the psychological profiles of large language models through psychometric inventories

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.961607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.961607Z digest=sha256:78858ec73979e76fabe1d0ea08321305d3e76315d8e4d28be492491e51e22379

Observation 0596c147-d6b9-4acf-b9a3-022ba8ba951a · outbound

This paper cites Discovering language model behaviors with model-written evaluations.

Value Drifts: Tracing Value Alignment During LLM Post-Training Discovering language model behaviors with model-written evaluations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.964623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.964623Z digest=sha256:6437b1bc2e7a26f9684b4ef3a1bf4563731af49446d02e43139f5ddb6a28a318

Observation 7a30bb37-6d7b-48c4-afa4-4b7135e157ea · outbound

This paper cites The lock-in hypothesis: Stagnation by algorithm.

Value Drifts: Tracing Value Alignment During LLM Post-Training The lock-in hypothesis: Stagnation by algorithm

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.967897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.967897Z digest=sha256:8a41d1d51bdbff0ad49fc77aae965c4114dd50a109ed7449036ba51370823353

Observation 1a50e33d-3126-4b41-93bd-cda1c32d5553 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Value Drifts: Tracing Value Alignment During LLM Post-Training Direct preference optimization: Your language model is secretly a reward model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.971048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.971048Z digest=sha256:075b3b3d1f56e840b4588978acdd575ca97e3e7a2e1f17e58296fd816fc23eeb

Observation 7ba371e8-31ca-488c-bd83-6c786723c494 · outbound

This paper cites Balancing the budget: Understanding trade-offs between supervised and preference-based finetuning.

Value Drifts: Tracing Value Alignment During LLM Post-Training Balancing the budget: Understanding trade-offs between supervised and preference-based finetuning

Reference 55

Resolution
verified exact
doi, observed 2026-08-04T07:23:25.417130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T07:21:36.974067Z digest=sha256:d76d0ad65985b0e9b35e08c2bd57fc852dc14cdf9b5df4acc4f251cecf4bb327

Observation c51d5a61-30a7-4562-8172-306c0c6917be · outbound

This paper cites Close Encounters of the AI Kind: A Survey of Public Sentiment About Artificial Intelligence.

Value Drifts: Tracing Value Alignment During LLM Post-Training Close Encounters of the AI Kind: A Survey of Public Sentiment About Artificial Intelligence

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.977018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.977018Z digest=sha256:5586355b74c065534de193e0aa6b4fa9c1602e9665dde9e5c024b46741601434

Observation bbe44da7-e415-4b51-972f-25ad4f82d755 · outbound

This paper cites Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them.

Value Drifts: Tracing Value Alignment During LLM Post-Training Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.979706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.979706Z digest=sha256:46c8f61fd14c6c7d00c8ab34819fbe28ba8c536bc220c2cf7278f9e162343f6a

Observation d757b45a-c349-4ce2-aaf2-3816fbd1fb33 · outbound

This paper cites Sentence- BERT : Sentence embeddings using S iamese BERT -networks.

Value Drifts: Tracing Value Alignment During LLM Post-Training Sentence- BERT : Sentence embeddings using S iamese BERT -networks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.982407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.982407Z digest=sha256:5135bd36b4bab02aa64d69cbdc04a457054f9fb46a69cbf5d7dc8c7f0fcf5f03

Observation 2b5d258e-ad5f-4136-b1f7-e69fb7b9f73f · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Value Drifts: Tracing Value Alignment During LLM Post-Training GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.985050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.985050Z digest=sha256:8caf4671428972a74159c346446c5f5a047b8be64863bb1e2490bccff1de3a61

Observation 9cf8c958-726a-4484-b5cf-3e8638269a19 · outbound

This paper cites Sutherland.

Value Drifts: Tracing Value Alignment During LLM Post-Training Sutherland

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.987575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.987575Z digest=sha256:f6b15127cf5b789a90feff26e93b01e6581b4305097ca557460e1d66941c9bea

Observation 09734dc1-b817-4710-9c0e-56e1f651171f · outbound

This paper cites The nature of human values.

Value Drifts: Tracing Value Alignment During LLM Post-Training The nature of human values

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.989983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.989983Z digest=sha256:77c2f9a096d68e09a8d342c40b3888fbd98ede100cc38cebe77b9516d1a3d970

Observation e5be2664-227d-4b88-b604-37a70899582b · outbound

This paper cites Political compass or spinning arrow? T owards more meaningful evaluations for values and opinions in large language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Political compass or spinning arrow? T owards more meaningful evaluations for values and opinions in large language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.992715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.992715Z digest=sha256:f4c67f8a8d8ba8f4f6a774071c074c3f0b913051111999432093536fcc67a61c

Observation 5800a519-118e-4268-95db-c0010128cf52 · outbound

This paper cites Unintended impacts of LLM alignment on global representation.

Value Drifts: Tracing Value Alignment During LLM Post-Training Unintended impacts of LLM alignment on global representation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.995122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.995122Z digest=sha256:9e8993246311952b77657aeaee8e9837eca4d174a13e5efbec0220320a4a76a0

Observation 4c2d37db-47a1-4192-ba6e-18113229c799 · outbound

This paper cites Personal values across cultures.

Value Drifts: Tracing Value Alignment During LLM Post-Training Personal values across cultures

Reference 64

Resolution
verified exact
doi, observed 2026-08-04T07:23:24.534884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T07:21:36.997512Z digest=sha256:da74b931a50365458f0a6663c64d76f7fe6b040c3b63a6a47e63d522648983d6

Observation e4e3e7f7-c2a8-47a6-873e-8c41f67f2a58 · outbound

This paper cites A note on the pure theory of consumer's behaviour.

Value Drifts: Tracing Value Alignment During LLM Post-Training A note on the pure theory of consumer's behaviour

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.000377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.000377Z digest=sha256:a4f581196243b21dfe4b049164a3adda3ca3693564407f98f9bd4ca11bf4ff0e

Observation 5e2dc7e3-fa82-40e3-b82b-a7a644cd1529 · outbound

This paper cites Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning, ICML'23.

Value Drifts: Tracing Value Alignment During LLM Post-Training Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning, ICML'23

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.003104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.003104Z digest=sha256:252f79d5cdc0fd61d700ec4fb403e0cb5179297895667db6fd6817dfe6a8914e

Observation c080ecf5-ddac-4e2b-a5bf-2dc02784bf35 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Value Drifts: Tracing Value Alignment During LLM Post-Training Proximal Policy Optimization Algorithms

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.006048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.006048Z digest=sha256:2edac0d7e895a45a3084304ec6bd7adabc356efda7664c99360592ef2b4ff145

Observation 6b244c73-9fa5-41c0-9394-57147c4ea694 · outbound

This paper cites Extending the cross-cultural validity of the theory of basic human values with a different method of measurement.

Value Drifts: Tracing Value Alignment During LLM Post-Training Extending the cross-cultural validity of the theory of basic human values with a different method of measurement

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.009189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.009189Z digest=sha256:bd71070e81cbce6d611c3f369f09ca1a3375c32658e2fbb54cc48925325e2f85

Observation 6b96b46b-896c-4566-8f79-cbc9bb3476f3 · outbound

This paper cites Personality Traits in Large Language Models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Personality Traits in Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.011514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.011514Z digest=sha256:d49dd8ae897c7e22718a4d33e1e3a05019c6c66da8e8a4611e8631bd3ab23a25

Observation 3644a10b-8e21-4e94-9b3f-f7117cf027b7 · outbound

This paper cites an unresolved cited work.

Value Drifts: Tracing Value Alignment During LLM Post-Training Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.014088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.014088Z digest=sha256:744c1ffc76f100bdb71416ef8175934573b37a4b55d470b7d15553cf41e3a789

Observation e959d165-fcc9-46a4-aacb-a4cab8e11f30 · outbound

This paper cites AI models collapse when trained on recursively generated data.

Value Drifts: Tracing Value Alignment During LLM Post-Training AI models collapse when trained on recursively generated data

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.016578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.016578Z digest=sha256:6c1889d0e6c9a9e739625136760df8d83b8b875f6ecf2a59597f0ff93700ab8d

Observation 6de127ee-bbd2-4f5b-8315-be8e6d85046a · outbound

This paper cites Recognizing stances in ideological on-line debates.

Value Drifts: Tracing Value Alignment During LLM Post-Training Recognizing stances in ideological on-line debates

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.018945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.018945Z digest=sha256:574f1ab528ed65de8b54afbaae3f9326c82bc1a2a406a44d1311b0a906adbdf1

Observation 75e6bea5-fa0f-497b-aed3-b9f49c0ce960 · outbound

This paper cites Position: A roadmap to pluralistic alignment.

Value Drifts: Tracing Value Alignment During LLM Post-Training Position: A roadmap to pluralistic alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.021453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.021453Z digest=sha256:4220ea85e56cfc31ada43c0c1b8600a88333d44fddcff13675b1cfff652ac734

Observation 598171ea-1410-4e46-a127-810c083de10a · outbound

This paper cites Value profiles for encoding human variation.

Value Drifts: Tracing Value Alignment During LLM Post-Training Value profiles for encoding human variation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.023833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.023833Z digest=sha256:f5dae66956fb9ad78ad4d60873ab7b1562633d5e92bc51dee47cfea4b02239d7

Observation 8e2756ac-3518-456d-91a9-8fb9e10c83cd · outbound

This paper cites Societal Alignment Frameworks Can Improve LLM Alignment.

Value Drifts: Tracing Value Alignment During LLM Post-Training Societal Alignment Frameworks Can Improve LLM Alignment

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.027536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.027536Z digest=sha256:0850ec40826531a4e249f642a3b1fae362befc89033f81a6b1b72ccaab96822f

Observation a3929b99-c3ad-4ab9-8b1c-5b98d9ffc56b · outbound

This paper cites Hashimoto.

Value Drifts: Tracing Value Alignment During LLM Post-Training Hashimoto

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.030262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.030262Z digest=sha256:862127550c6cff65f261dd29b61a5227395b4016503c7e35ac26b245a58966e9

Observation 525a74ce-e610-4566-b478-a56051b46833 · outbound

This paper cites A deep dive into the trade-offs of parameter-efficient preference alignment techniques.

Value Drifts: Tracing Value Alignment During LLM Post-Training A deep dive into the trade-offs of parameter-efficient preference alignment techniques

Reference 77

Resolution
verified exact
doi, observed 2026-08-04T07:23:23.815623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-04T07:21:37.032655Z digest=sha256:9b67f7ec1db7babc07d2d75dba9c1b7ffd9a71519103ef41b79ec599cfe1d0c3

Observation d9b4f115-1e60-4c8c-a6da-3861ad1d6efd · outbound

This paper cites Zephyr: Direct distillation of LM alignment.

Value Drifts: Tracing Value Alignment During LLM Post-Training Zephyr: Direct distillation of LM alignment

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.035278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.035278Z digest=sha256:d53b5eeb31b6e3d28ec62fde95dd52d0f51a824b6728c90dd6e38a6489eac4b7

Observation 042852bf-7a90-444e-a523-934f7e9a50ab · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

Value Drifts: Tracing Value Alignment During LLM Post-Training Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.037596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.037596Z digest=sha256:268d8b7896605e4bbcb1b37c0c2ffac9d410d50416c14bbed3af9f0b584cb02f

Observation 3b4fe45e-3de7-43ec-9337-9acab483fec0 · outbound

This paper cites Dai, and Quoc V Le.

Value Drifts: Tracing Value Alignment During LLM Post-Training Dai, and Quoc V Le

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.040157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.040157Z digest=sha256:57cef279ccd72379c4fdeb342fd19f3b24ca994b85186192772f7cc28344b3ee

Observation c6df399c-3182-45d0-8f2c-5b66a7ad10d5 · outbound

This paper cites Simple synthetic data reduces sycophancy in large language models, 2025.

Value Drifts: Tracing Value Alignment During LLM Post-Training Simple synthetic data reduces sycophancy in large language models, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.042675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.042675Z digest=sha256:64f12bb6ff45026ca4d046cd90a0a4503ce580b0ce2e7720ee0f7fc5b1ac1ea3

Observation a1f2541f-3de0-4329-920f-ca7f7e606755 · outbound

This paper cites Generative monoculture in large language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Generative monoculture in large language models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.045646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.045646Z digest=sha256:c8ad003c1f741066ac286ba473885331b0fe3710c78aa075cb507e721637b0c0

Observation 903165b2-286b-4b5c-a49d-430bcbd09f15 · outbound

This paper cites Fairness feedback loops: T raining on synthetic data amplifies bias.

Value Drifts: Tracing Value Alignment During LLM Post-Training Fairness feedback loops: T raining on synthetic data amplifies bias

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.048330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.048330Z digest=sha256:09c9c43bad65396c6c13a37f3376687e2e125f06e8de7cb0f7de80dea49fb358

Observation 2859d098-492b-4a6e-b637-604ab0131a74 · outbound

This paper cites Finding the sweet spot: Preference data construction for scaling preference optimization.

Value Drifts: Tracing Value Alignment During LLM Post-Training Finding the sweet spot: Preference data construction for scaling preference optimization

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.051409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.051409Z digest=sha256:ae23f13c82c6ebac5cf64d906483eead72d838a3f62a197a5f47d8effc26735e

Observation 57f17ede-74b7-42fe-8765-87296f6ee870 · outbound

This paper cites Qwen3 Technical Report.

Value Drifts: Tracing Value Alignment During LLM Post-Training Qwen3 Technical Report

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.053895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.053895Z digest=sha256:9befa26dc79201898e3d36439e937b63a74757f2f3a989c7502a2922f587c9c8

Observation 243287a8-3e34-4a96-956e-826044370593 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Value Drifts: Tracing Value Alignment During LLM Post-Training HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.056559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.056559Z digest=sha256:c24339dca0c59349ffbfe90f1b1ffd770fa928771527cbfdc959e304d5219936

Observation bda9dd5c-5a7f-424b-8526-9546765e2a12 · outbound

This paper cites Cultivating pluralism in algorithmic monoculture: The community alignment dataset.

Value Drifts: Tracing Value Alignment During LLM Post-Training Cultivating pluralism in algorithmic monoculture: The community alignment dataset

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.059357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.059357Z digest=sha256:6d7617f2b999282b404760eae0d1f170f37ea0e90c998971e6e01f9a760dbe7b

Observation 822b0dd1-0a95-4955-87e4-6c232c4c4d42 · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Value Drifts: Tracing Value Alignment During LLM Post-Training Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.061716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.061716Z digest=sha256:fdb47dc19ab34f5466303964c28ea9abd7309dc0e1bd60057bfaa3668e45b0b2

Observation 7bdb39b3-4714-428e-bd87-d71216056727 · outbound

This paper cites WildChat : 1m chat GPT interaction logs in the wild.

Value Drifts: Tracing Value Alignment During LLM Post-Training WildChat : 1m chat GPT interaction logs in the wild

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.064385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.064385Z digest=sha256:6f421664a2139fafc307df11504b37710a96d5411504354d6b0f2ccfa175b825

Observation 79a032b9-b362-4938-b476-819ec8889aea · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Value Drifts: Tracing Value Alignment During LLM Post-Training Secrets of RLHF in Large Language Models Part I: PPO

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.066762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.066762Z digest=sha256:e59e53bb6e803eaa4fcddcfb4b3341d686e64c4535c1ed7f6d5ddf6399b9a36b

Observation 45df0724-a0af-40ba-a0bd-b59f1653f31a · outbound

This paper cites @esa (Ref.

Value Drifts: Tracing Value Alignment During LLM Post-Training @esa (Ref

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.069555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.069555Z digest=sha256:71f6c3b8ef4e0074a8601240697134c3f745a256a67f4ebdc06aaeff811a6f21

Observation 5ee8303d-dbe4-40e0-94fb-0da5de1c300d · outbound

This paper cites an unresolved cited work.

Value Drifts: Tracing Value Alignment During LLM Post-Training Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.072580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.072580Z digest=sha256:d9a7fe27571af93a0dc22900ad5201cb6add65fb032a25475cb8d3e1c604c688

Observation ae830681-f92a-46a2-9b68-618736407641 · outbound

This paper cites small value-gap.

Value Drifts: Tracing Value Alignment During LLM Post-Training small value-gap

Reference 93

Resolution
malformed identifier
no resolver link, observed 2026-08-04T07:21:37.075273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.075273Z digest=sha256:c5d41d45db7dfcc51b396d2bf135159626f45d9874aa84d887dc1c8def80c2dc

Pith citing papers

Observation 2b36d67e-7fce-43c3-aaf8-ea4c70cb0075 · inbound

Agents of Chaos cites this paper.

Agents of Chaos Value Drifts: Tracing Value Alignment During LLM Post-Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-16T02:21:43.072236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T07:03:41.003127Z digest=sha256:f60fd518cba1735bc253de0b01980ecba8ffb918413057b7ba2eaad9c1f2d658

Observation 1dc20ae4-13db-441a-9fc1-a700325ee9cf · inbound

Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editing cites this paper.

Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editing Value Drifts: Tracing Value Alignment During LLM Post-Training

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-16T02:21:43.072236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-03T14:25:39.401131Z digest=sha256:91e4e41b4bc65f44d5fdffeecec3ee59aab0b41d428a4a5eca08b66bde21111e