Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:44:12.908742Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2607.22676.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T07:44:12.908742Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 15c20bf2-6b42-4332-97ca-365fa69fed23 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b9ed988-c8cd-455e-86de-0ffd23a790bc · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4db81838-fe98-44f2-b42a-6cf38c5448e7 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Shashwat Goel, Rishi Hazra, Dulhan Jayalath, Timon Willi, Parag Jain, William F Shen, Ilias Leontiadis, Francesco Barbieri, Yoram Bachrach, Jonas Geiping, et al
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1d13e2b-96d8-4e02-8bde-7704aebf7b2f · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Towards an AI co-scientist
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 893a902e-5c71-448f-b6ab-b9601a943054 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1c8b897-c2c1-455b-bbcd-af557fd70f26 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Alignment faking in large language models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b686f350-5d9c-45fd-b691-3b5b48c10978 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67bbd316-71c9-4e0c-aa08-37a9e75ccdb4 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a95bba1c-e718-4851-bb67-d551f67aa950 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Val-bench: Measuring value alignment in language models.arXiv preprint arXiv:2510.05465,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60817586-e035-49c2-9647-8334b689781a · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift What is in Your Safe Data? Identifying Benign Data that Breaks Safety
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93edbb92-ec61-45e5-a8d3-97889d1de0cf · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Understanding catastrophic forgetting in language models via implicit inference
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed4e70e1-67fb-424a-937f-79b556eadb57 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd3fdea8-e9d5-448b-8413-30ca16bb3482 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Holistic Evaluation of Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa646ea2-3a36-4082-bb39-13ac93b65480 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f1373da-6470-4aa8-8236-455916a01d57 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Natural emergent misalignment from reward hacking in production RL.arXiv preprint arXiv:2511.18397,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac092efd-0bf9-41b7-8ef4-d2f87a8cea56 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da97eac8-1fb3-4ae9-8c6b-0d23d3ed6c0e · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18ae449e-e1b3-4d26-bef0-54597ca2b9c6 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Steering Llama 2 via Contrastive Activation Addition
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a7d4085-942d-4811-a70d-16f853adca04 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift BBQ: A Hand-Built Bias Benchmark for Question Answering
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e611791f-350a-4d55-914d-eda50aacda53 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Evaluating Frontier Models for Stealth and Situational Awareness
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a909cb6d-3143-4ced-a1df-19e32ac9708d · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d90204a-e197-4ddb-829f-a2dc575a1c93 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Proximal Policy Optimization Algorithms
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faa928a4-7c31-4015-a6e7-8f38383d3c1d · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Towards understanding sycophancy in language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9799853d-1213-49ec-9fed-4a72e5596838 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift SEAT: Sparse Entity-Aware Tuning for Knowledge Adaptation while Preserving Epistemic Abstention
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecc72661-25ad-4704-bebd-5a06674b1d0d · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Rethinking rubric generation for improving llm judge and reward modeling for open-ended tasks.arXiv preprint arXiv:2602.05125, 2026b
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4383b94-2b9f-4d14-9b4a-b5ffc5bd0505 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Efficiency vs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a576681-4e36-4855-a189-130589f46210 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ccb2f31-6291-4be4-81d2-34523d9259e8 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44ce3f33-e5bb-444d-b68f-21fc76dd0456 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Continual Learning for Large Language Models: A Survey
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c758f0d3-8683-4048-894a-f186754125ba · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Qwen2.5 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3282e5e5-ac55-46e1-aca2-a48f7d1b8956 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Qwen3 Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dc0aa6b-640a-4bd9-8fbc-554c13c04373 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bc464dc-398e-45a6-a3c0-388d06036cc7 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Removing RLHF Protections in GPT-4 via Fine-Tuning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af3ebd5a-7c02-4ae3-8603-1ef177ca39df · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Instruction-Following Evaluation for Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ea0e6d-4bc8-479c-8011-a47dc483ee69 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift The path not taken: RLVR provably learns off the principals.arXiv preprint arXiv:2511.08567,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd97763e-1c09-4b9f-824d-79d1451e595e · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Representation Engineering: A Top-Down Approach to AI Transparency
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bb0f0b0-23d6-40c2-a88b-c47c2bfe9f17 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift 15 A.2 KL-Regularized SFT
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29f6243c-f362-4ae0-8bd9-ab00a9b7110f · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift GRPO-based RLVR is trained to a fixed number of steps with almost all reaching reward saturation, while SFT and KL-SFT use the fixed epoch budgets in Table
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21c96e9d-c54a-45a6-873c-e0280e23026d · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 240f51fc-e227-4227-86ff-10f2289e7486 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift We report the headline metric, the preferred direction, and the aggregation procedure used when benchmarks contain multiple subtasks
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9305019-ea83-44a8-a90a-13cdd5fe1d29 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47ed0b9f-43c9-4d4a-b30e-33b65240793b · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Unresolved cited work
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53feb5b4-32ca-43f4-b1cf-38789ac341f3 · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Breaking the safety-capability tradeoff: Reinforcement learning with verifiable rewards maintains safety guardrails in LLMs.arXiv preprint arXiv:2511.21050,
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f170256-bcb4-45cb-b0a5-042892823d5f · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift TACO: Topics in Algorithmic COde generation dataset
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 715f6f3d-a357-440f-8bbd-1efd2c690acf · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2b3e20b-dbc5-4ab0-8f0f-77258d1ed2fb · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Program Synthesis with Large Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be94e655-cd29-4999-9bef-8b9fb268711a · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Evaluating Large Language Models Trained on Code
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca54fa19-a4f1-410c-986d-48769cb5dcee · outbound
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift Persona features control emergent misalignment.arXiv preprint arXiv:2506.19823,
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.