Pith. sign in

Paper Citation Record · LEDGER

Value Drifts: Tracing Value Alignment During LLM Post-Training

As of 20 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 3 inbound Pith citation observations for arXiv:2510.26707.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.26707 v2

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:21:37.075273Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:25:16.255780Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:28:31.100259Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved89
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b7020dab-58e3-4a4e-8b27-e7ddb9b3a545 · outbound

This paper cites write newline.

Value Drifts: Tracing Value Alignment During LLM Post-Training write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.162650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.162650Z digest=sha256:531bf731f52b0fb5e86f29042b7b1ad7c533759eae1962a5e6f3b4aa565aae25

Observation 583ba2b3-030e-489b-b61a-edf3aaf2545b · outbound

This paper cites Llama 3 model card.

Value Drifts: Tracing Value Alignment During LLM Post-Training Llama 3 model card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.254871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.254871Z digest=sha256:31214f506de58191020cf67eb3d738f99bdb350c346ee920ef4ba8447f00ca0b

Observation c691439b-6c02-4607-97ce-bce0e5225742 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.304935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.304935Z digest=sha256:1b6fbb5c29fc438871ff04bf63f1efaf2bcd6ec2a39603293a5ad02962c35d7f

Observation ccea4413-468e-4c0d-a1d5-00ca6ef01197 · outbound

This paper cites Explicitly unbiased large language models still form biased associations.

Value Drifts: Tracing Value Alignment During LLM Post-Training Explicitly unbiased large language models still form biased associations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.415062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.415062Z digest=sha256:b0386dd800f71bf3440221404f09fefb5afa9e53146a26457878a170e9f82582

Observation 25ef445c-80f4-4d22-b188-e1e7d5403e31 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Value Drifts: Tracing Value Alignment During LLM Post-Training Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.505290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.505290Z digest=sha256:8b02a993ebb11b1f7e9689322cde06de3765809b46d347431e88f49a80eb5796

Observation 868e3cb9-8791-4c68-bf8d-ce68f25f16a6 · outbound

This paper cites Managing extreme AI risks amid rapid progress.

Value Drifts: Tracing Value Alignment During LLM Post-Training Managing extreme AI risks amid rapid progress

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.644736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.644736Z digest=sha256:3e6f5c716a1e39f7502ca2616f8af54e6115e5e8ded9ab27f827492c23e8d3ca

Observation bec17b05-1907-4ff2-b9d9-eef95268efad · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

Value Drifts: Tracing Value Alignment During LLM Post-Training PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.759175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.759175Z digest=sha256:f28b9437c74e96b20e493cb1fa5d5c04b6debd334f08c6eea06268de89a3724c

Observation 7d92a9b5-5d41-4c2a-aaad-d14ed3ff6f7e · outbound

This paper cites Picking on the same person: Does algorithmic monoculture lead to outcome homogenization? Advances in Neural Information Processing Systems, 35: 0 3663--3678, 2022.

Value Drifts: Tracing Value Alignment During LLM Post-Training Picking on the same person: Does algorithmic monoculture lead to outcome homogenization? Advances in Neural Information Processing Systems, 35: 0 3663--3678, 2022

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.848184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.848184Z digest=sha256:097868f69be795822e7ac1a960ac849516880905ee175b9d389b64885f97689d

Observation d25f0d5f-e33a-479a-89d8-bd8935ca3324 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Value Drifts: Tracing Value Alignment During LLM Post-Training Rank analysis of incomplete block designs: I

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.895674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.895674Z digest=sha256:5b74d3a973dd05d593f20c3f1a0d6cb3919bacd51cb8f5793222a0ffd0be57f7

Observation ff12c554-9e1a-449e-b44b-ca34341a3352 · outbound

This paper cites Density-based clustering based on hierarchical density estimates.

Value Drifts: Tracing Value Alignment During LLM Post-Training Density-based clustering based on hierarchical density estimates

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:32.989406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:32.989406Z digest=sha256:8fcf8f04766eae317df01d7f7a078a3c63e54464b42a42bdc67998e7455dea5e

Observation 61c25588-6734-4e04-a503-e6ac06b173d4 · outbound

This paper cites How people use chatgpt.

Value Drifts: Tracing Value Alignment During LLM Post-Training How people use chatgpt

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:33.151838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:33.151838Z digest=sha256:4efb552121f57c49eff452bac8e824344ab9ed30223a126e10f22cbc3d7b2e1f

Observation f8af7be2-4772-4162-95cd-7d2cd92e1e12 · outbound

This paper cites Chatbot arena: An open platform for evaluating LLMs by human preference.

Value Drifts: Tracing Value Alignment During LLM Post-Training Chatbot arena: An open platform for evaluating LLMs by human preference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:33.275598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:33.275598Z digest=sha256:1d5f7931fed6a88336cf0b51cf7f52bddffd3643c1490fd927201deecba1591e

Observation 80661ef1-2db9-42e3-a4c4-78ed71680965 · outbound

This paper cites Reward model interpretability via optimal and pessimal tokens.

Value Drifts: Tracing Value Alignment During LLM Post-Training Reward model interpretability via optimal and pessimal tokens

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:33.428012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:33.428012Z digest=sha256:b74a2dd523fb25676cb62bfb44ec419f21feb09d8cacffc4bc602806a292b111

Observation c0854b5b-b415-4338-9639-f691b53da6dd · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2023.

Value Drifts: Tracing Value Alignment During LLM Post-Training Ultrafeedback: Boosting language models with high-quality feedback, 2023

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:33.594774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:33.594774Z digest=sha256:9dc1dcee7f990e0d07ab1a44d7e33d0807f68f392b2fe28886cdd6971a48acdf

Observation 8e58ee50-3c08-4035-bb6a-a2d358894d86 · outbound

This paper cites Towards measuring the representation of subjective global opinions in language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Towards measuring the representation of subjective global opinions in language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:33.744615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:33.744615Z digest=sha256:289a2024606d2a3369f85ae3bfd1de4ba678f7b06d106629010389bdd7e42304

Observation ff4fe65e-d1f5-4eb6-96dc-0c36871e3cd9 · outbound

This paper cites Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective.

Value Drifts: Tracing Value Alignment During LLM Post-Training Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:33.898254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:33.898254Z digest=sha256:b645fdec155f5cd88a59fed9ed7fa0241ac657de80555152b4dbe78919e318c3

Observation 4e89affd-d33a-40ea-ad19-cdf217241346 · outbound

This paper cites Artificial intelligence, values, and alignment.

Value Drifts: Tracing Value Alignment During LLM Post-Training Artificial intelligence, values, and alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:34.060310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:34.060310Z digest=sha256:9bf760742a07ed6ba14f543cc82d97a6f33aa9297e7dffa74e6b34e699460175

Observation 4e5dfb86-0977-4182-ab54-2e041776e0e3 · outbound

This paper cites The delta learning hypothesis: Preference tuning on weak data can yield strong gains.

Value Drifts: Tracing Value Alignment During LLM Post-Training The delta learning hypothesis: Preference tuning on weak data can yield strong gains

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:34.179117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:34.179117Z digest=sha256:91b76fbc0e6ad71c3c610eb461701d3680e6c530da81a1e20f774769cdf3b86e

Observation 8db42bf4-9729-41d2-b6d9-c1eba768a891 · outbound

This paper cites Donoho, and Sanmi Koyejo.

Value Drifts: Tracing Value Alignment During LLM Post-Training Donoho, and Sanmi Koyejo

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:34.326465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:34.326465Z digest=sha256:47c5805180784f76894f040439c126baece6653afc7b1569fe5d819a3e9ecadd

Observation 11313824-2d8c-4dac-84da-31195cb06a85 · outbound

This paper cites Mixed Preference Optimization: Reinforcement Learning with Data Selection and Better Reference Model.

Value Drifts: Tracing Value Alignment During LLM Post-Training Mixed Preference Optimization: Reinforcement Learning with Data Selection and Better Reference Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:34.487518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:34.487518Z digest=sha256:f06127ac5c6846c7e2be83c6bda36ad6e3983bdaaf1c11cb87355ed0819d6912

Observation c969a66e-9580-4f57-b164-ca65bdc3adca · outbound

This paper cites Alignment faking in large language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Alignment faking in large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:34.645305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:34.645305Z digest=sha256:0e935bef5e61253051024ca8921fbf36986c3ef8bcfa48abe3c49fff94428846

Observation db5c1e40-539c-49b6-9dc1-a79331402d02 · outbound

This paper cites Assessing the alignment of large language models with human values for mental health integration: Cross-sectional study using schwartz’s theory of basic values.

Value Drifts: Tracing Value Alignment During LLM Post-Training Assessing the alignment of large language models with human values for mental health integration: Cross-sectional study using schwartz’s theory of basic values

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:34.807216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:34.807216Z digest=sha256:f1f859686c7480958e84dc8ca7cf37b0cee43162b07643a123777d97f1fcb6d6

Observation 7081315d-ac38-409d-8dee-b50f99117000 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Value Drifts: Tracing Value Alignment During LLM Post-Training Measuring Massive Multitask Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:34.925316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:34.925316Z digest=sha256:b209a0b6afa38fff44561cc2d2bbdbd4b748a5d906579255a6630beb9b485459

Observation 4df6632a-359b-49ae-94ca-c2ce52d9cf17 · outbound

This paper cites Collective constitutional AI : Aligning a language model with public input.

Value Drifts: Tracing Value Alignment During LLM Post-Training Collective constitutional AI : Aligning a language model with public input

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.048765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.048765Z digest=sha256:9e576349dbb817c0ad89ee75c9c7587b62e7caeba7b2aec55915dbab180d959b

Observation c59fcb09-fbd9-4b40-8cd6-a72464373a79 · outbound

This paper cites Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions.

Value Drifts: Tracing Value Alignment During LLM Post-Training Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.154191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.154191Z digest=sha256:80668f44d6e0ee27d307568b2d54702388ab3a99cdb286e2e40cfe929d1e570d

Observation f7171a3d-7da3-43e6-992e-c3a88d358e18 · outbound

This paper cites The n+ implementation details of RLHF with PPO : A case study on TL ; DR summarization.

Value Drifts: Tracing Value Alignment During LLM Post-Training The n+ implementation details of RLHF with PPO : A case study on TL ; DR summarization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.236114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.236114Z digest=sha256:ad9138e78d3401f0b89d45bd49042b8c53a198590c118cda74d5156f8fd19989

Observation 7d834588-3883-4372-ace6-2a3ea5455fb9 · outbound

This paper cites Smith, Yejin Choi, and Hannaneh Hajishirzi.

Value Drifts: Tracing Value Alignment During LLM Post-Training Smith, Yejin Choi, and Hannaneh Hajishirzi

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.313021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.313021Z digest=sha256:3b3296f9c67dc7626660a4b264078ae77ccbcf9009522df74a42e8bb2bf01f87

Observation 43fa4ae1-344f-42cd-92a8-5ae510799db5 · outbound

This paper cites Evaluating and inducing personality in pre-trained language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Evaluating and inducing personality in pre-trained language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.393478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.393478Z digest=sha256:0063f7541975489b658597ea46a070223198642521266738e06ef509d26393b3

Observation 6664693e-d4f1-4378-b89a-bfb1aa9d12ff · outbound

This paper cites Can Machines Learn Morality? The Delphi Experiment.

Value Drifts: Tracing Value Alignment During LLM Post-Training Can Machines Learn Morality? The Delphi Experiment

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.508909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.508909Z digest=sha256:3c2bec2db5ddab1a8030a37a6ddc5df861fd4738a87405929397589700d561f3

Observation 8c019ce7-207a-4777-93f6-ce1786575ea5 · outbound

This paper cites The PRISM alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training The PRISM alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.671360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.671360Z digest=sha256:3d18b9e049963690e56d6eb369be7307a1c98735122ba5e3bf54c8042888c544

Observation e4f0ae8d-e82a-4240-b65c-68f0c77ce503 · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Value Drifts: Tracing Value Alignment During LLM Post-Training Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:35.867471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:35.867471Z digest=sha256:5e853e065a861898901ef9b77e8dd36a474e5776a45128a289cfc97cb7b88cfe

Observation 799a5af8-7389-4679-a38b-7c4e8bf47bf9 · outbound

This paper cites What are human values, and how do we align AI to them?.

Value Drifts: Tracing Value Alignment During LLM Post-Training What are human values, and how do we align AI to them?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.038024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.038024Z digest=sha256:66aa6e0ff760d5fd1a8fc68758d080335c79b311e57bcc79dc7e3581b80e8615

Observation 9b3ed33b-b11d-49c6-9727-d10fa1c087e9 · outbound

This paper cites You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation.

Value Drifts: Tracing Value Alignment During LLM Post-Training You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.204883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.204883Z digest=sha256:cf56155faffa9703baea36c090cca757ceb7fc1c67cabcc276003a440b94669c

Observation 6a5cdd90-821f-44c8-aea4-7e8a7dcc8ae5 · outbound

This paper cites Beyond probabilities: Unveiling the misalignment in evaluating large language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Beyond probabilities: Unveiling the misalignment in evaluating large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.365788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.365788Z digest=sha256:37f9d9754c6aa3f73693c3a9b31d3dfce258bb6f868898d54984907cf432ab25

Observation c1c841ae-ac45-4cfc-9560-50c6fcbf4233 · outbound

This paper cites Treleaven, and Miguel Rodrigues Rodrigues.

Value Drifts: Tracing Value Alignment During LLM Post-Training Treleaven, and Miguel Rodrigues Rodrigues

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.527874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.527874Z digest=sha256:459d7b397b42c40f5d20bf90acc5e2a37fd7f6ad073c641e909ca0f663350ba2

Observation 401a553f-a0ce-4630-9c9a-6d94c8cc87f6 · outbound

This paper cites How people use claude for support, advice, and companionship, 2025.

Value Drifts: Tracing Value Alignment During LLM Post-Training How people use claude for support, advice, and companionship, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.652253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.652253Z digest=sha256:e4257562130e7ef123cf728c8a8e2158c1a11ab11b18406d8df8e808f8097717

Observation e0443012-cc9c-43b4-85a3-a507ff5a1f90 · outbound

This paper cites UMAP : Uniform manifold approximation and projection.

Value Drifts: Tracing Value Alignment During LLM Post-Training UMAP : Uniform manifold approximation and projection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.772987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.772987Z digest=sha256:de35356a7d2ad15225ae2451b54abf774324c3830540a99876e5e00a46449090

Observation c41dd467-5975-4b76-9836-c30c62f78fff · outbound

This paper cites Sim PO : Simple preference optimization with a reference-free reward.

Value Drifts: Tracing Value Alignment During LLM Post-Training Sim PO : Simple preference optimization with a reference-free reward

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.887916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.887916Z digest=sha256:4357011ffea3a73ad298ee92216510490fb9b28a2cdfe133082341e748cc2299

Observation ee67d119-5e3b-4ce7-bdec-3780c02cec85 · outbound

This paper cites S em E val-2016 task 6: Detecting stance in tweets.

Value Drifts: Tracing Value Alignment During LLM Post-Training S em E val-2016 task 6: Detecting stance in tweets

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.923689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.923689Z digest=sha256:e7d643cb63655701e2ca07d5cd5422d1211fb217911ecb1a4ec069a41b35d841

Observation 7c69da23-1635-43ab-a2f5-28564340eecd · outbound

This paper cites an unresolved cited work.

Value Drifts: Tracing Value Alignment During LLM Post-Training Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.927567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.927567Z digest=sha256:c23ed69740772a5d615b81c69b899b7cc0a011fa549e14b434f0cc113b0e8f10

Observation d68db222-c491-47f3-966b-20f2c0522de4 · outbound

This paper cites Reinforcement learning finetunes small subnetworks in large language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Reinforcement learning finetunes small subnetworks in large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.930799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.930799Z digest=sha256:98650fd6bec2ee881e9f415934028de3805fcd913c4d6fe56b0a239adaaa11a7

Observation 4e02c4ea-5139-49a9-ad6c-d877154124cb · outbound

This paper cites Value imprint: A technique for auditing the human values embedded in RLHF datasets.

Value Drifts: Tracing Value Alignment During LLM Post-Training Value imprint: A technique for auditing the human values embedded in RLHF datasets

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.933807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.933807Z digest=sha256:ade28c4f13933419d05be9e6a7e9f332d61e2c82a6c0afd8fe98f5c177f6a681

Observation 4506d2a1-8ed9-4f8f-a201-446923a8bc83 · outbound

This paper cites Attributing mode collapse in the fine-tuning of large language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Attributing mode collapse in the fine-tuning of large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.937221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.937221Z digest=sha256:ea0c812a06b4c37fee77a6c83a762887176f0a3dfb28b8c52612f3e6874a4b2a

Observation 93f02863-4f38-4f68-b74f-dcc50402cecf · outbound

This paper cites Help OpenAI fix over-refusals! https://community.openai.com/t/help-openai-fix-over-refusals/409799, October 2023.

Value Drifts: Tracing Value Alignment During LLM Post-Training Help OpenAI fix over-refusals! https://community.openai.com/t/help-openai-fix-over-refusals/409799, October 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.940405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.940405Z digest=sha256:4b156f77d9b203643163ee38d0c3b52c5aecfe9fb78ca70bd6ef4b351e5c9223

Observation cb844eee-afd8-4989-819d-406d159fb918 · outbound

This paper cites Training language models to follow instructions with human feedback.

Value Drifts: Tracing Value Alignment During LLM Post-Training Training language models to follow instructions with human feedback

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.943222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.943222Z digest=sha256:b5f99ce80b220d6a560c2cebbde73c4e217fe6a1bf5547ecc6b0d08da40b2624

Observation 2163ede4-d52a-41e7-8815-17c97486c2d1 · outbound

This paper cites Does Writing with Language Models Reduce Content Diversity?.

Value Drifts: Tracing Value Alignment During LLM Post-Training Does Writing with Language Models Reduce Content Diversity?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.946254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.946254Z digest=sha256:913994dd77cd0cb463c4e2eaf40272b57ae1ff0e1dcf939fe834d34ac731f6f5

Observation 7612a28b-e51f-4648-9590-bcdca58dfb72 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Value Drifts: Tracing Value Alignment During LLM Post-Training Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.949344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.949344Z digest=sha256:38ccdcee6df5a798272b345c2659bbe6c64543076d3f129daf32beb5a3440575

Observation f487b967-7c27-41da-acc1-d30af7ec929b · outbound

This paper cites Do LLMs Possess a Personality? Making the MBTI Test an Amazing Evaluation for Large Language Models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Do LLMs Possess a Personality? Making the MBTI Test an Amazing Evaluation for Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.952385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.952385Z digest=sha256:bd745c17b72c0b9245820b78a89db9211785b011a459fdbbee2b4c5087972834

Observation b4505e6d-be81-4c75-b378-f56b154576a8 · outbound

This paper cites What matters in data for DPO ? arXiv preprint arXiv:2508.18312, 2025.

Value Drifts: Tracing Value Alignment During LLM Post-Training What matters in data for DPO ? arXiv preprint arXiv:2508.18312, 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.955565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.955565Z digest=sha256:d001032d8a3ac9a7ec1a090ee12fcbe18f5fa5cf688c4dccf189e32096972385

Observation a30fd34e-c19f-41c8-bfba-41909370aeb9 · outbound

This paper cites Enhancing alignment using curriculum learning & ranked preferences.

Value Drifts: Tracing Value Alignment During LLM Post-Training Enhancing alignment using curriculum learning & ranked preferences

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.958465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.958465Z digest=sha256:c51cd381b05e4f384e8f4181351fd826435422ee7de0ae04fcb33d68fe6086b6

Observation 8b0332f1-9254-4451-9e14-ac841cfceef4 · outbound

This paper cites AI psychometrics: Assessing the psychological profiles of large language models through psychometric inventories.

Value Drifts: Tracing Value Alignment During LLM Post-Training AI psychometrics: Assessing the psychological profiles of large language models through psychometric inventories

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.961607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.961607Z digest=sha256:ee653cb9d7bc6705b14ca303130ee2a9e4f8c572a0174d15836dcf7ac79098fa

Observation 0596c147-d6b9-4acf-b9a3-022ba8ba951a · outbound

This paper cites Discovering language model behaviors with model-written evaluations.

Value Drifts: Tracing Value Alignment During LLM Post-Training Discovering language model behaviors with model-written evaluations

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.964623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.964623Z digest=sha256:d95c963d9ba057b00289eddab8036e732b0a2712de0b3491fa98bfb84df29ded

Observation 7a30bb37-6d7b-48c4-afa4-4b7135e157ea · outbound

This paper cites The lock-in hypothesis: Stagnation by algorithm.

Value Drifts: Tracing Value Alignment During LLM Post-Training The lock-in hypothesis: Stagnation by algorithm

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.967897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.967897Z digest=sha256:f00d4cb5d93121351ed84e5c9976d43e705ab847039e4b470568bf4c0a0267d1

Observation 1a50e33d-3126-4b41-93bd-cda1c32d5553 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Value Drifts: Tracing Value Alignment During LLM Post-Training Direct preference optimization: Your language model is secretly a reward model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.971048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.971048Z digest=sha256:46ff59ea247c23e8a1dc9ddc49aa07b87dad88818fa159099890bc6805135ed7

Observation 7ba371e8-31ca-488c-bd83-6c786723c494 · outbound

This paper cites Balancing the budget: Understanding trade-offs between supervised and preference-based finetuning.

Value Drifts: Tracing Value Alignment During LLM Post-Training Balancing the budget: Understanding trade-offs between supervised and preference-based finetuning

Reference 55

Resolution
verified exact
doi, observed 2026-08-04T07:23:25.417130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-04T07:21:36.974067Z digest=sha256:ca32b282a650481f24b7579227d9e93ffb616c56a79cddc08fa4c10c6e54991e

Observation c51d5a61-30a7-4562-8172-306c0c6917be · outbound

This paper cites Close Encounters of the AI Kind: A Survey of Public Sentiment About Artificial Intelligence.

Value Drifts: Tracing Value Alignment During LLM Post-Training Close Encounters of the AI Kind: A Survey of Public Sentiment About Artificial Intelligence

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.977018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.977018Z digest=sha256:94e80b206b2d0016c5a1398f3fac5bbf45e5ce1e07bbe899a14e147773f59667

Observation bbe44da7-e415-4b51-972f-25ad4f82d755 · outbound

This paper cites Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them.

Value Drifts: Tracing Value Alignment During LLM Post-Training Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.979706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.979706Z digest=sha256:5ffeed60f3c46867c7bc3f940441d578f1892c7131398e36856c8523ec0ff857

Observation d757b45a-c349-4ce2-aaf2-3816fbd1fb33 · outbound

This paper cites Sentence- BERT : Sentence embeddings using S iamese BERT -networks.

Value Drifts: Tracing Value Alignment During LLM Post-Training Sentence- BERT : Sentence embeddings using S iamese BERT -networks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.982407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.982407Z digest=sha256:3bfb57a8110413f43f62220b6d415669d1c48aa3b2c327ca83801bd5a5bd8ad5

Observation 2b5d258e-ad5f-4136-b1f7-e69fb7b9f73f · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Value Drifts: Tracing Value Alignment During LLM Post-Training GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.985050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.985050Z digest=sha256:a5c6c7da34ba8455741ba739470fcf2adfc15300fcb1d5dab6def697d51300ba

Observation 9cf8c958-726a-4484-b5cf-3e8638269a19 · outbound

This paper cites Sutherland.

Value Drifts: Tracing Value Alignment During LLM Post-Training Sutherland

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.987575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.987575Z digest=sha256:dcb08ec26797c1de9a989768d0879a1a12aa0f6e44bfc6565efe76da43524dc3

Observation 09734dc1-b817-4710-9c0e-56e1f651171f · outbound

This paper cites The nature of human values.

Value Drifts: Tracing Value Alignment During LLM Post-Training The nature of human values

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.989983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.989983Z digest=sha256:0d4f27b9bc62f878b19ce5f6356a219e50625c7a36cfa41c05544cfb0bf8459d

Observation e5be2664-227d-4b88-b604-37a70899582b · outbound

This paper cites Political compass or spinning arrow? T owards more meaningful evaluations for values and opinions in large language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Political compass or spinning arrow? T owards more meaningful evaluations for values and opinions in large language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.992715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.992715Z digest=sha256:03fba15c82e02fbbdf19b48db16b454879a235ef6d2cdae1f3eb152a0ce98c34

Observation 5800a519-118e-4268-95db-c0010128cf52 · outbound

This paper cites Unintended impacts of LLM alignment on global representation.

Value Drifts: Tracing Value Alignment During LLM Post-Training Unintended impacts of LLM alignment on global representation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:36.995122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:36.995122Z digest=sha256:6c4b92b3cd41210947f3172b0eab39bed56e79f4e9d053d81b214d25df7b696c

Observation 4c2d37db-47a1-4192-ba6e-18113229c799 · outbound

This paper cites Personal values across cultures.

Value Drifts: Tracing Value Alignment During LLM Post-Training Personal values across cultures

Reference 64

Resolution
verified exact
doi, observed 2026-08-04T07:23:24.534884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-04T07:21:36.997512Z digest=sha256:451fd32f6f759ab9bf567c46d7901d1b8c6f52cbbb6aa1ba43b239a3126cecc2

Observation e4e3e7f7-c2a8-47a6-873e-8c41f67f2a58 · outbound

This paper cites A note on the pure theory of consumer's behaviour.

Value Drifts: Tracing Value Alignment During LLM Post-Training A note on the pure theory of consumer's behaviour

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.000377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.000377Z digest=sha256:686e6a7f960228019903d276611399a6ad7e26ddb9734c04b56403a68c5c08ec

Observation 5e2dc7e3-fa82-40e3-b82b-a7a644cd1529 · outbound

This paper cites Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning, ICML'23.

Value Drifts: Tracing Value Alignment During LLM Post-Training Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning, ICML'23

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.003104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.003104Z digest=sha256:b81701ba4c6349d505f959a32eab9441dff4e8e310ac22e1a0a74943e37c016f

Observation c080ecf5-ddac-4e2b-a5bf-2dc02784bf35 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Value Drifts: Tracing Value Alignment During LLM Post-Training Proximal Policy Optimization Algorithms

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.006048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.006048Z digest=sha256:91dddab77b53bf5479fe77730f482953862aaee018546d74a7871d1cfc0258c2

Observation 6b244c73-9fa5-41c0-9394-57147c4ea694 · outbound

This paper cites Extending the cross-cultural validity of the theory of basic human values with a different method of measurement.

Value Drifts: Tracing Value Alignment During LLM Post-Training Extending the cross-cultural validity of the theory of basic human values with a different method of measurement

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.009189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.009189Z digest=sha256:f1f3e3902e3f95fc3d92c10152480caccb795c49d3c54c6eb07ba3f2326994f5

Observation 6b96b46b-896c-4566-8f79-cbc9bb3476f3 · outbound

This paper cites Personality Traits in Large Language Models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Personality Traits in Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.011514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.011514Z digest=sha256:b823c858dedeb0b10c83284aff8b82bdf5fd2f1f22a29430dc7c75b77a3ee186

Observation 3644a10b-8e21-4e94-9b3f-f7117cf027b7 · outbound

This paper cites an unresolved cited work.

Value Drifts: Tracing Value Alignment During LLM Post-Training Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.014088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.014088Z digest=sha256:d1be032f49584516aa18f9db2fb3f5d1f946cc771dfbbaf22c245e2bf5083cfa

Observation e959d165-fcc9-46a4-aacb-a4cab8e11f30 · outbound

This paper cites AI models collapse when trained on recursively generated data.

Value Drifts: Tracing Value Alignment During LLM Post-Training AI models collapse when trained on recursively generated data

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.016578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.016578Z digest=sha256:7e49a6da724de3cf7267afc993d2940295ed9808a350bfd3b4b2f3d24c22e917

Observation 6de127ee-bbd2-4f5b-8315-be8e6d85046a · outbound

This paper cites Recognizing stances in ideological on-line debates.

Value Drifts: Tracing Value Alignment During LLM Post-Training Recognizing stances in ideological on-line debates

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.018945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.018945Z digest=sha256:9582274bf7f8e90d761c663bdfeb514fbc9f97fcf6602f55fb9e0f3ecef837f7

Observation 75e6bea5-fa0f-497b-aed3-b9f49c0ce960 · outbound

This paper cites Position: A roadmap to pluralistic alignment.

Value Drifts: Tracing Value Alignment During LLM Post-Training Position: A roadmap to pluralistic alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.021453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.021453Z digest=sha256:c8029e9a1c2ed4dd046f44696cf8aa7f812ed76a35915eda31362ec4d5b00d83

Observation 598171ea-1410-4e46-a127-810c083de10a · outbound

This paper cites Value profiles for encoding human variation.

Value Drifts: Tracing Value Alignment During LLM Post-Training Value profiles for encoding human variation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.023833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.023833Z digest=sha256:fd9355f7b063cc728d44f176e4992abc0d0abba444302a19964d16dc2b7dac99

Observation 8e2756ac-3518-456d-91a9-8fb9e10c83cd · outbound

This paper cites Societal Alignment Frameworks Can Improve LLM Alignment.

Value Drifts: Tracing Value Alignment During LLM Post-Training Societal Alignment Frameworks Can Improve LLM Alignment

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.027536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.027536Z digest=sha256:257c247e6d2d547da57888723979a07778b5158d9229d1d34d9f79ce650d947a

Observation a3929b99-c3ad-4ab9-8b1c-5b98d9ffc56b · outbound

This paper cites Hashimoto.

Value Drifts: Tracing Value Alignment During LLM Post-Training Hashimoto

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.030262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.030262Z digest=sha256:408aaaff5804b7bc94329cadc11253b11464f7cd1420e9097f37ed3502cdf5b1

Observation 525a74ce-e610-4566-b478-a56051b46833 · outbound

This paper cites A deep dive into the trade-offs of parameter-efficient preference alignment techniques.

Value Drifts: Tracing Value Alignment During LLM Post-Training A deep dive into the trade-offs of parameter-efficient preference alignment techniques

Reference 77

Resolution
verified exact
doi, observed 2026-08-04T07:23:23.815623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-04T07:21:37.032655Z digest=sha256:f4946bef4a5b9372aa7e6c623f78e9cadeb081d068e9b50d72e452cfab11558a

Observation d9b4f115-1e60-4c8c-a6da-3861ad1d6efd · outbound

This paper cites Zephyr: Direct distillation of LM alignment.

Value Drifts: Tracing Value Alignment During LLM Post-Training Zephyr: Direct distillation of LM alignment

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.035278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.035278Z digest=sha256:480ae3a1bd7d90ce8960d4d8541d83265f622403c431d7046aecf2854f5a9932

Observation 042852bf-7a90-444e-a523-934f7e9a50ab · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

Value Drifts: Tracing Value Alignment During LLM Post-Training Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.037596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.037596Z digest=sha256:9f26f2f5ec7b8ccf01519959dc0d3abe135643806fa704c878fcc8f944d27e50

Observation 3b4fe45e-3de7-43ec-9337-9acab483fec0 · outbound

This paper cites Dai, and Quoc V Le.

Value Drifts: Tracing Value Alignment During LLM Post-Training Dai, and Quoc V Le

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.040157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.040157Z digest=sha256:8e7431e6c5df83b4b5acdba4866f043e4cca19e80e713b35e3bbc55cbe62030e

Observation c6df399c-3182-45d0-8f2c-5b66a7ad10d5 · outbound

This paper cites Simple synthetic data reduces sycophancy in large language models, 2025.

Value Drifts: Tracing Value Alignment During LLM Post-Training Simple synthetic data reduces sycophancy in large language models, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.042675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.042675Z digest=sha256:12f983eb816899b7025051b52732ce803e6ccc3a60d027f778e3b18bff59a451

Observation a1f2541f-3de0-4329-920f-ca7f7e606755 · outbound

This paper cites Generative monoculture in large language models.

Value Drifts: Tracing Value Alignment During LLM Post-Training Generative monoculture in large language models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.045646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.045646Z digest=sha256:b82a6ee164bc8537dc4c474ee09c3dbc9e7ef1a471dac97d1110c7a71e3cda1a

Observation 903165b2-286b-4b5c-a49d-430bcbd09f15 · outbound

This paper cites Fairness feedback loops: T raining on synthetic data amplifies bias.

Value Drifts: Tracing Value Alignment During LLM Post-Training Fairness feedback loops: T raining on synthetic data amplifies bias

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.048330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.048330Z digest=sha256:471464e124beef48184c131cd99addca1924af879101181183e25eae5236d2e1

Observation 2859d098-492b-4a6e-b637-604ab0131a74 · outbound

This paper cites Finding the sweet spot: Preference data construction for scaling preference optimization.

Value Drifts: Tracing Value Alignment During LLM Post-Training Finding the sweet spot: Preference data construction for scaling preference optimization

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.051409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.051409Z digest=sha256:182576a62448c663794dd65998accd49e1478b7cb9a6669532261a10b727cd0e

Observation 57f17ede-74b7-42fe-8765-87296f6ee870 · outbound

This paper cites Qwen3 Technical Report.

Value Drifts: Tracing Value Alignment During LLM Post-Training Qwen3 Technical Report

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.053895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.053895Z digest=sha256:f5e5448cb488604dfd451845bc42ecf8d77bd982efef77b5e90c4adff5076fa5

Observation 243287a8-3e34-4a96-956e-826044370593 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Value Drifts: Tracing Value Alignment During LLM Post-Training HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.056559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.056559Z digest=sha256:ee872486411e2f11a4ec7a35d60dbdf83fa7ef6bfafec988c6bcd97c343504bd

Observation bda9dd5c-5a7f-424b-8526-9546765e2a12 · outbound

This paper cites Cultivating pluralism in algorithmic monoculture: The community alignment dataset.

Value Drifts: Tracing Value Alignment During LLM Post-Training Cultivating pluralism in algorithmic monoculture: The community alignment dataset

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.059357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.059357Z digest=sha256:3cb1f2af6dcdb47fa5fa9b7540993ce39f610f8b13de18499544495c282dda77

Observation 822b0dd1-0a95-4955-87e4-6c232c4c4d42 · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Value Drifts: Tracing Value Alignment During LLM Post-Training Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.061716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.061716Z digest=sha256:70316879f7534e5bf49893cc56e932a8ab2855b5e444ba6b265165232dc3168f

Observation 7bdb39b3-4714-428e-bd87-d71216056727 · outbound

This paper cites WildChat : 1m chat GPT interaction logs in the wild.

Value Drifts: Tracing Value Alignment During LLM Post-Training WildChat : 1m chat GPT interaction logs in the wild

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.064385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.064385Z digest=sha256:6ad95d28337ba8ce66e234597fad3750f700ae3f36f70f32d56f2650485707e1

Observation 79a032b9-b362-4938-b476-819ec8889aea · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Value Drifts: Tracing Value Alignment During LLM Post-Training Secrets of RLHF in Large Language Models Part I: PPO

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.066762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.066762Z digest=sha256:d18ffa229cf3d3939fa02417f87e77a7be85c51b6ead657a34dcd5852151d69f

Observation 45df0724-a0af-40ba-a0bd-b59f1653f31a · outbound

This paper cites @esa (Ref.

Value Drifts: Tracing Value Alignment During LLM Post-Training @esa (Ref

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.069555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.069555Z digest=sha256:89a0bced1e45b1035bd21047682498ab541f63cc95b723608d13ff18ef11a60e

Observation 5ee8303d-dbe4-40e0-94fb-0da5de1c300d · outbound

This paper cites an unresolved cited work.

Value Drifts: Tracing Value Alignment During LLM Post-Training Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.072580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.072580Z digest=sha256:37f89d5a28ab38cde30f568a33c8e3dc730af6cee26931526259f4e8564cd6e9

Observation ae830681-f92a-46a2-9b68-618736407641 · outbound

This paper cites small value-gap.

Value Drifts: Tracing Value Alignment During LLM Post-Training small value-gap

Reference 93

Resolution
malformed identifier
no resolver link, observed 2026-08-04T07:21:37.075273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.075273Z digest=sha256:f71ccc514a4fb132075149c1bdac91b559eee661e38b782877821f4dc0024158

Pith citing papers

Observation 2b36d67e-7fce-43c3-aaf8-ea4c70cb0075 · inbound

Agents of Chaos cites this paper.

Agents of Chaos Value Drifts: Tracing Value Alignment During LLM Post-Training

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-16T02:21:43.072236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T07:03:41.003127Z digest=sha256:d4bc1159193d64b2350cc00421f50396f071fbf4575d413357451038f21b51d1

Observation 1dc20ae4-13db-441a-9fc1-a700325ee9cf · inbound

Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editing cites this paper.

Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editing Value Drifts: Tracing Value Alignment During LLM Post-Training

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-16T02:21:43.072236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-03T14:25:39.401131Z digest=sha256:e3903518780b692c8e83ede3bb3e888a695b32b975bb28b039035d376e9045a5

Observation f7bdd74a-b723-4e27-af8a-0aeddb49edca · inbound

Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases cites this paper.

Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases Value Drifts: Tracing Value Alignment During LLM Post-Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:25:16.255780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:25:16.255780Z digest=sha256:50a1fb598101d1a8d3a289e9f1bdb994ca9db960d1e5975320a71674792cfd5b