Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:58.258118Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2509.25148.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:58.258118Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T00:48:56.892634Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T21:18:58.373323Z
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f85890b8-4fab-4908-b811-2d4220c38e01 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Harnessing the power of llms in practice: A survey on chatgpt and beyond
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7e4e84ab-65ee-43a4-9e33-dee86a050071 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9048366-9d1e-4557-b7db-57d92e51a4dc · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models \texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3e67ee99-bf16-4aeb-b499-937ff5aa3806 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models A survey on post-training of large language models.arXiv e-prints, pages arXiv–2503, 2025
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4fdda5c9-fb6f-4973-82ea-31fa4c32784f · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Training language models to follow instructions with human feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f8a0d64-a97c-4051-83e3-f8f4f95bf819 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Deep reinforcement learning from human preferences
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d98b61c7-d0b4-4444-8999-b0f0948a7d50 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Direct preference optimization: Your language model is secretly a reward model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8aec489-bfc7-4e0d-9795-013e3887e282 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56686d20-3a5d-490d-bb73-6df25066f822 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Curriculum offline imitating learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 11144b59-b2cf-4452-a379-532cd0493a36 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Offline imitation learning with suboptimal demonstrations via relaxed distribution matching
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e38fa1cb-b58c-4008-acaa-fc7a3123f3b3 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Offline Reinforcement Learning for LLM Multi-Step Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cd45998-58e6-4d1a-b03b-681575012957 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Grounding large language models in interactive environments with online reinforcement learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a34bb80f-a6e5-47ef-b26c-5de71dbef4dc · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Provably mitigating overoptimization in rlhf: Your sft loss is implicitly an adversarial regularizer
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 12dcd533-12d3-4efd-8ae2-85c5e8b1fe88 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Preserving Diversity in Supervised Fine-Tuning of Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 070ec0c1-259d-4d4b-8a41-30d2908456ab · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Regularizing Neural Networks by Penalizing Confident Output Distributions
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b801e931-70b3-47d3-9822-e7b9710433c2 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Stanford alpaca: An instruction-following llama model, 2023
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2b49f18d-da45-43a6-b2be-234497283dd4 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2316bf61-624d-421a-83bc-f0cf68cc97f7 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models On the diversity of synthetic data and its impact on training large language models,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b8b1e8c8-d5b7-413e-9e0c-7d90f7b60170 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Condor: Enhance llm alignment with knowledge-driven data synthesis and refinement
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 938e0ef4-961e-4a9c-9218-bd996afa1a9b · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Idgen: Item discrimination induced prompt generation for llm evaluation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d27af825-ab83-4627-be9a-6228d708508d · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Qwen3 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08e3b7df-4394-48c2-85f5-199226c1483e · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DeepSeek-V3 Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fcddce0-5c59-4725-bca2-b15566824e6e · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Gemma: Open Models Based on Gemini Research and Technology
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98c886ce-718d-4454-8c6e-ccafe91ea064 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models GPT-4 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23866afd-1047-4d4b-b9ec-a3a32e138b27 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d439eeeb-0b36-4daa-8288-a0f1779a44bd · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df333a4f-9aea-4408-801c-1d20fc8e213d · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c7ec438-faeb-4e0b-ba3b-1a95d3c8c691 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Reinforcement learning with verifiable rewards: Grpo’s effective loss, dynam- ics, and success amplification
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 028c0ad8-5942-4d92-9c1f-aaa47d435aab · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models AutoGLM: Autonomous Foundation Agents for GUIs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ebdcc37-7fe3-49c2-b0a6-2d60a82ad4c6 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7faba00-97d3-403c-a438-df0d71050587 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04c91798-d270-49d5-8e48-98e58360054b · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98ef0e3a-b85d-4bb8-9e5d-8a6ec1c0af87 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models On-policy rl meets off-policy experts: Harmonizing supervised fine- tuning and reinforcement learning via dynamic weighting
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3ecce42-6552-4fbe-b36c-1906c1d830e9 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Learning to Reason under Off-Policy Guidance
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8858a5b0-372f-42e1-a199-4c9d910cbcb6 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5c519e1-9d20-4eb9-af14-6b057420aa87 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 875a2f1d-728a-487a-90bc-606c0c1d27da · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a198c5bc-aa3d-440b-bb9a-8d46304380f5 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing, 2024
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54df15d6-af5d-4049-993c-1bde54bbe74c · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Generalizing verifiable instruction following, 2025
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a762c342-dd82-4265-8ff4-5eedfda017f5 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Instruction-Following Evaluation for Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92da0776-98c2-4f65-8985-74d5e47917c4 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f85c6f8f-bac3-47e1-b658-f4e938b68f5b · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Measuring Massive Multitask Language Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ee6b06f-52d5-4c6c-b909-a4c1087c866c · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 804d1d4c-4045-4555-a32d-91e48d1258b3 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Gpqa: A graduate-level google-proof q&a benchmark
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b947f471-b0f1-400f-af49-1deb956d0934 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Evaluating Large Language Models Trained on Code
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0979e746-9fab-4609-beee-71afeb3e1c52 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Program Synthesis with Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6005055c-4816-4fd3-97ff-ab9624d64ae7 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Training Verifiers to Solve Math Word Problems
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7ace46a-de22-4611-9396-116de9b0e6f1 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Let’s verify step by step
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82d5c9da-f8f5-4afa-8ba3-be80aee45ae7 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Theoremqa: A theorem-driven question answering dataset
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 62e6de16-ca27-4c2d-9a22-65beb8e7d1d8 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Cmmlu: Measuring massive multitask language understanding in chinese, 2023
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8fcaa28-bd44-49ef-b07e-4760195a12b7 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models C-eval: 12 A multi-level multi-discipline chinese evaluation suite for foundation models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ebb3d37d-99fb-4248-8f04-9148819eb483 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Pre-trained policy discriminators are general reward models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 324a8893-47d2-4c6a-a407-5dbb8abfa522 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Swift:a scalable lightweight infrastructure for fine-tuning, 2025
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ac153536-d7aa-48c5-9f68-870b934ff8ad · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Hybridflow: A flexible and efficient rlhf framework
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6522c83-15cd-4dc4-8a16-7f7d405c00b9 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models preference
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5c6a21b2-369d-4393-95a4-8a8e50d8b392 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b4ef34fe-e198-4a5d-950e-33621b0407d7 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 35491afa-7d8e-49b3-b244-848509f433bb · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Due to the properties of the Dirac delta function, the integral is non-zero only at the single point y=y ∗
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a999e21e-f008-46c7-bd41-fb96a7093c29 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 64913ba6-5505-4e22-8bbf-9bb3998ba328 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9bd36c44-6590-42cb-863f-e34b69f0c504 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a35edc2e-255f-43c8-a8bc-2e8c9c63e1c1 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models This provides a cohesive narrative for the entire post-training pipeline, viewing it not as a sequence of disparate steps but as a unified process
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d01692d6-8b7b-4e18-935f-1834b976261b · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models winner" (preferred) response andyl is the
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 829c96dc-3465-4610-bb3e-b7248a0a1296 · outbound
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models On the Diversity of Synthetic Data and its Impact on Training Large Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82e3a9ec-f72b-42da-b747-44ef9f62ab08 · inbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.