Pith. sign in

Paper Citation Record · LEDGER

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models

As of 16 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2509.25148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.25148 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:58.258118Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T00:48:56.892634Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T21:18:58.373323Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved47
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f85890b8-4fab-4908-b811-2d4220c38e01 · outbound

This paper cites Harnessing the power of llms in practice: A survey on chatgpt and beyond.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Harnessing the power of llms in practice: A survey on chatgpt and beyond

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.792700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.061817Z digest=sha256:76c0a3aac61ed20f41f2dad7e4c898f5546659ca94ad566e0299b238a69e5c86

Observation 7e4e84ab-65ee-43a4-9e33-dee86a050071 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.065373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.065373Z digest=sha256:a180744051d345d32d9b3876bd8e84e65750befae6acf7f28dfa19f17303914a

Observation d9048366-9d1e-4557-b7db-57d92e51a4dc · outbound

This paper cites \texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models \texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T15:49:58.605246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.068948Z digest=sha256:b8254208bbf6b92760f0da5ad3c82e2d9b9125757b2f8095087a8b00fc6a5d78

Observation 3e67ee99-bf16-4aeb-b499-937ff5aa3806 · outbound

This paper cites A survey on post-training of large language models.arXiv e-prints, pages arXiv–2503, 2025.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models A survey on post-training of large language models.arXiv e-prints, pages arXiv–2503, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.785712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.072181Z digest=sha256:4e087f51a314ae59529f8f59af0fb8511fa3473243a87c34dc05e807c38d07f5

Observation 4fdda5c9-fb6f-4973-82ea-31fa4c32784f · outbound

This paper cites Training language models to follow instructions with human feedback.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Training language models to follow instructions with human feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.075991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.075991Z digest=sha256:0973c5eb13aa3f6f74d9abb18e1cce2d994c5c78d2291cf5671e72c671838934

Observation 4f8a0d64-a97c-4051-83e3-f8f4f95bf819 · outbound

This paper cites Deep reinforcement learning from human preferences.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.078529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.078529Z digest=sha256:43148a753e8ade4df5451315bdf46e50cc04108719396a2d7d96c221235d5f0d

Observation d98b61c7-d0b4-4444-8999-b0f0948a7d50 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.081745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.081745Z digest=sha256:965698e4553a5f648b548d873f0a631efbbb972df97df293d8d918ba0fabc89a

Observation c8aec489-bfc7-4e0d-9795-013e3887e282 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.084732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.084732Z digest=sha256:be0d7e41ecf8f4ad49be895c089084b6a9ab38d56dacb4cb191c9f32b2a70565

Observation 56686d20-3a5d-490d-bb73-6df25066f822 · outbound

This paper cites Curriculum offline imitating learning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Curriculum offline imitating learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.768174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.088393Z digest=sha256:30c070fd541b2fc89144bd7259de322002442039fed60cc848508304aa4ca118

Observation 11144b59-b2cf-4452-a379-532cd0493a36 · outbound

This paper cites Offline imitation learning with suboptimal demonstrations via relaxed distribution matching.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Offline imitation learning with suboptimal demonstrations via relaxed distribution matching

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.761477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.091500Z digest=sha256:8c9610a04602f36c0057cd278ab4f6d90211a16a62369b939b19fa2817974eb8

Observation e38fa1cb-b58c-4008-acaa-fc7a3123f3b3 · outbound

This paper cites Offline Reinforcement Learning for LLM Multi-Step Reasoning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.094264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.094264Z digest=sha256:f398bd1c94791b71408a3b6633b369ef2865563780bae607197e8f5e386dc83f

Observation 1cd45998-58e6-4d1a-b03b-681575012957 · outbound

This paper cites Grounding large language models in interactive environments with online reinforcement learning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Grounding large language models in interactive environments with online reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.097622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.097622Z digest=sha256:3bb89843dbe39715c1d22ba41826ec507a9beb0c7eabb32fc849f92e0bc9372a

Observation a34bb80f-a6e5-47ef-b26c-5de71dbef4dc · outbound

This paper cites Provably mitigating overoptimization in rlhf: Your sft loss is implicitly an adversarial regularizer.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Provably mitigating overoptimization in rlhf: Your sft loss is implicitly an adversarial regularizer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.751138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.100869Z digest=sha256:8afe1d873f1d552ad10adb115806a7f1acbe80d74270d42eecc1d02eb56e89a2

Observation 12dcd533-12d3-4efd-8ae2-85c5e8b1fe88 · outbound

This paper cites Preserving Diversity in Supervised Fine-Tuning of Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Preserving Diversity in Supervised Fine-Tuning of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.103306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.103306Z digest=sha256:f0a5913a69b3ec74f114ea9c4bb73fdbb4d731f44ede8bc98850f38f7df4f44b

Observation 070ec0c1-259d-4d4b-8a41-30d2908456ab · outbound

This paper cites Regularizing Neural Networks by Penalizing Confident Output Distributions.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Regularizing Neural Networks by Penalizing Confident Output Distributions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.106522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.106522Z digest=sha256:584652faad12a0f01243796587e6d1e7ea0fc03acbd9a8aa80c376cb8b7b2f6d

Observation b801e931-70b3-47d3-9822-e7b9710433c2 · outbound

This paper cites Stanford alpaca: An instruction-following llama model, 2023.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Stanford alpaca: An instruction-following llama model, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.743340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.109290Z digest=sha256:e7cfcff0cc151aac97220eb1c23662922a586729428823fe46d777e631d34fac

Observation 2b49f18d-da45-43a6-b2be-234497283dd4 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.111627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.111627Z digest=sha256:b43c407ef547993a9a9c98c5b58b01f53a83e5a1ef7c93acdf7159cdc283863b

Observation 2316bf61-624d-421a-83bc-f0cf68cc97f7 · outbound

This paper cites On the diversity of synthetic data and its impact on training large language models,.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models On the diversity of synthetic data and its impact on training large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.736182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.114383Z digest=sha256:dfe877b0b40c5ca0ac73fcf3450b44228ac45e56a390ad8543fca21b0dc9b154

Observation b8b1e8c8-d5b7-413e-9e0c-7d90f7b60170 · outbound

This paper cites Condor: Enhance llm alignment with knowledge-driven data synthesis and refinement.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Condor: Enhance llm alignment with knowledge-driven data synthesis and refinement

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.728110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.119357Z digest=sha256:fd6d94044ff672e2cd67f6db89d3ec75528fb1677694a14bcb92ee07aff5ce16

Observation 938e0ef4-961e-4a9c-9218-bd996afa1a9b · outbound

This paper cites Idgen: Item discrimination induced prompt generation for llm evaluation.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Idgen: Item discrimination induced prompt generation for llm evaluation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.721234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.123053Z digest=sha256:5dfcbbb8a60e9c99da7ffe68382b4a83dc21365e823541f0df57f2801a3fa219

Observation d27af825-ab83-4627-be9a-6228d708508d · outbound

This paper cites Qwen3 Technical Report.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Qwen3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.126327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.126327Z digest=sha256:84384c495f0983aea71d684f30369c611edf9f7b88b45c985826fa00b0ffdc3d

Observation 08e3b7df-4394-48c2-85f5-199226c1483e · outbound

This paper cites DeepSeek-V3 Technical Report.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DeepSeek-V3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.129064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.129064Z digest=sha256:fc2a4e129c2c45e904db8aef46b47aa4cc40c583962f78b93079a07a5c400731

Observation 4fcddce0-5c59-4725-bca2-b15566824e6e · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.132125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.132125Z digest=sha256:7a36ea25f37f0adbb84ac4f4ea41f91b9567dab3df92b31996d35adebc6d7e50

Observation 98c886ce-718d-4454-8c6e-ccafe91ea064 · outbound

This paper cites GPT-4 Technical Report.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.136271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.136271Z digest=sha256:649a60ed96e74c318b3e0a7619e543c07061a3ad281ed09d62efc1aaa508aa00

Observation 23866afd-1047-4d4b-b9ec-a3a32e138b27 · outbound

This paper cites Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.138798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.138798Z digest=sha256:bf1ed88152377a6b6285be515ac99f0189819cb3ecaf5b4bae7a446b1e9ad358

Observation d439eeeb-0b36-4daa-8288-a0f1779a44bd · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.142255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.142255Z digest=sha256:a20a78072843ec6bba10779e08eaa86f8523ef64388db78b450df80418b12a2c

Observation df333a4f-9aea-4408-801c-1d20fc8e213d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.146538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.146538Z digest=sha256:eadcaa3f7b6dd0f3e339e3424356ca4969b0a89764cc20ccb0dc2cbfe5453148

Observation 7c7ec438-faeb-4e0b-ba3b-1a95d3c8c691 · outbound

This paper cites Reinforcement learning with verifiable rewards: Grpo’s effective loss, dynam- ics, and success amplification.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Reinforcement learning with verifiable rewards: Grpo’s effective loss, dynam- ics, and success amplification

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.150276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.150276Z digest=sha256:4432bdd9b7fa7b4504f42ed71ae58a32f6630fc104f02e07cfc22091d93a2433

Observation 028c0ad8-5942-4d92-9c1f-aaa47d435aab · outbound

This paper cites AutoGLM: Autonomous Foundation Agents for GUIs.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models AutoGLM: Autonomous Foundation Agents for GUIs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.152922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.152922Z digest=sha256:a5f759adc742821eb0f0aebce4a8411acdf470607d9236ce206302c38349014a

Observation 7ebdcc37-7fe3-49c2-b0a6-2d60a82ad4c6 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.155753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.155753Z digest=sha256:6213ef10173cf8e274521e844f1f6f7d9d7572d2025e18cff5ef8197b27a39a2

Observation a7faba00-97d3-403c-a438-df0d71050587 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.158558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.158558Z digest=sha256:0068e032e06145851814cd1e6171613d80c2910e55c8bbecc557cd63ca79ce62

Observation 04c91798-d270-49d5-8e48-98e58360054b · outbound

This paper cites Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.162168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.162168Z digest=sha256:63d420bcab7fc95f218de348fdf51078bd4b569cf07fef0a507b7379f97ccba5

Observation 98ef0e3a-b85d-4bb8-9e5d-8a6ec1c0af87 · outbound

This paper cites On-policy rl meets off-policy experts: Harmonizing supervised fine- tuning and reinforcement learning via dynamic weighting.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models On-policy rl meets off-policy experts: Harmonizing supervised fine- tuning and reinforcement learning via dynamic weighting

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.165261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.165261Z digest=sha256:1530c4895de256527f8c07c15905630d1eea44e9025ac2c8af1ec8ceab8e8154

Observation b3ecce42-6552-4fbe-b36c-1906c1d830e9 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Learning to Reason under Off-Policy Guidance

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.169948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.169948Z digest=sha256:d990991ef179f30a5d7d68dc804f3020948fdccbbd551fe81025e1585bcee475

Observation 8858a5b0-372f-42e1-a199-4c9d910cbcb6 · outbound

This paper cites BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.173572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.173572Z digest=sha256:e05e73bf825217755c23ef3a673e477f8f300011ee13f40682d629baed8bdded

Observation a5c519e1-9d20-4eb9-af14-6b057420aa87 · outbound

This paper cites SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.176602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.176602Z digest=sha256:8a30135c69532572bf5b59a8dc8230593bc614f33c0ae75aadfc1c76a4d4c8d7

Observation 875a2f1d-728a-487a-90bc-606c0c1d27da · outbound

This paper cites Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.179658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.179658Z digest=sha256:6eb64e6eaee61948e1f5cb92a38298011c75a5f8f7eac2b838f932791c329022

Observation a198c5bc-aa3d-440b-bb9a-8d46304380f5 · outbound

This paper cites Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing, 2024.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.183190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.183190Z digest=sha256:7cba5f68079ac4a03b311b7764149c85024cce65361fae65556441c5aef76ffa

Observation 54df15d6-af5d-4049-993c-1bde54bbe74c · outbound

This paper cites Generalizing verifiable instruction following, 2025.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Generalizing verifiable instruction following, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.185865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.185865Z digest=sha256:4e452a920c0d7139c43fd8c485ab3d1bc61b91534f34cc188c303fe1a56139b0

Observation a762c342-dd82-4265-8ff4-5eedfda017f5 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Instruction-Following Evaluation for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.188396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.188396Z digest=sha256:310f1081b10014161110667200645de6f9fdd83530d7ce6055c0d87baef829be

Observation 92da0776-98c2-4f65-8985-74d5e47917c4 · outbound

This paper cites Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.190607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.190607Z digest=sha256:623722ed9b7904f3968dea6aba7385b8d3dc0426ede05b8bf14c6fc652e16bc9

Observation f85c6f8f-bac3-47e1-b658-f4e938b68f5b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Measuring Massive Multitask Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.193411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.193411Z digest=sha256:a3aaf5247abe44dd48d2c34cb3bd4f1b687853a6182df98bf2970660317882f3

Observation 9ee6b06f-52d5-4c6c-b909-a4c1087c866c · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.196925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.196925Z digest=sha256:a2d9115aab1ebd3cd537d944cbd4928f332e800e6f980afd33928dd554c2523d

Observation 804d1d4c-4045-4555-a32d-91e48d1258b3 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Gpqa: A graduate-level google-proof q&a benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.199271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.199271Z digest=sha256:c881b2030e4214a0b4a1712d5852e94853053bb93d4740e8c5a26f105ec81925

Observation b947f471-b0f1-400f-af49-1deb956d0934 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Evaluating Large Language Models Trained on Code

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.201429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.201429Z digest=sha256:f5268cc4cfb565b4a68e2d16729b4531d06c8f431bced967f2b6c030da2be404

Observation 0979e746-9fab-4609-beee-71afeb3e1c52 · outbound

This paper cites Program Synthesis with Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Program Synthesis with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.205422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.205422Z digest=sha256:fa218fc43a7cf86e768ae144fb2653f486d52a1cb45f558d4ee03bca33bd94d5

Observation 6005055c-4816-4fd3-97ff-ab9624d64ae7 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.209138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.209138Z digest=sha256:3311e78da5f5c6645a584d2cdcece95e56c782f2b54e04b20a451a8d7a656ff5

Observation b7ace46a-de22-4611-9396-116de9b0e6f1 · outbound

This paper cites Let’s verify step by step.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Let’s verify step by step

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.212019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.212019Z digest=sha256:6efaa12eacff6d4a6fa59b23264559bf4fb035a8c0f2d37170bb0b56ed24d715

Observation 82d5c9da-f8f5-4afa-8ba3-be80aee45ae7 · outbound

This paper cites Theoremqa: A theorem-driven question answering dataset.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Theoremqa: A theorem-driven question answering dataset

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.699052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.214895Z digest=sha256:5ea6dad8d24441929dfdfddc0301c11e25a07f6a094b5fa7a3a88e4b254bacbb

Observation 62e6de16-ca27-4c2d-9a22-65beb8e7d1d8 · outbound

This paper cites Cmmlu: Measuring massive multitask language understanding in chinese, 2023.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Cmmlu: Measuring massive multitask language understanding in chinese, 2023

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.217412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.217412Z digest=sha256:18d90b77783989d9b50ddc760ec5845e5ffd40b20443783b109f48336d62dac0

Observation e8fcaa28-bd44-49ef-b07e-4760195a12b7 · outbound

This paper cites C-eval: 12 A multi-level multi-discipline chinese evaluation suite for foundation models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models C-eval: 12 A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.688924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.219886Z digest=sha256:4eb71142f51a30c135fb01f07a316f48118aa39324e4802941add1fed7b83e98

Observation ebb3d37d-99fb-4248-8f04-9148819eb483 · outbound

This paper cites Pre-trained policy discriminators are general reward models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Pre-trained policy discriminators are general reward models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.222765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.222765Z digest=sha256:49f38578bb6d1a021a25ad516301cbd55934a15c5ff702cd7cb41416db83c12f

Observation 324a8893-47d2-4c6a-a407-5dbb8abfa522 · outbound

This paper cites Swift:a scalable lightweight infrastructure for fine-tuning, 2025.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Swift:a scalable lightweight infrastructure for fine-tuning, 2025

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.681698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.225967Z digest=sha256:4bebbc5a30c81eeba20d7c585b48c3ac0b04e0f14aa9fe77a29fb19650ead63b

Observation ac153536-d7aa-48c5-9f68-870b934ff8ad · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Hybridflow: A flexible and efficient rlhf framework

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.228336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.228336Z digest=sha256:3986f11e2e6adbb04dd75c3bc7b16f17abf9e1c74886db789dd04c8fc94a594c

Observation f6522c83-15cd-4dc4-8a16-7f7d405c00b9 · outbound

This paper cites preference.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models preference

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.671519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.230938Z digest=sha256:93a31eac91179fa61da6bd0b9303992a03f311fd8a62451bee1497982e105215

Observation 5c6a21b2-369d-4393-95a4-8a8e50d8b392 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.663350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.234960Z digest=sha256:df1ff3abb184c26dd4c2170ef925d49593a0eb607abc886710dc97a4d22818c6

Observation b4ef34fe-e198-4a5d-950e-33621b0407d7 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.655942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.238143Z digest=sha256:af86b0f20ef87a94f0e2d67d182ed4fc0503aa45b664afd59c50ddf9ec35ed39

Observation 35491afa-7d8e-49b3-b244-848509f433bb · outbound

This paper cites Due to the properties of the Dirac delta function, the integral is non-zero only at the single point y=y ∗.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Due to the properties of the Dirac delta function, the integral is non-zero only at the single point y=y ∗

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.646748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.240961Z digest=sha256:c1af13de6878c3a608789249a0d966d41e808af0b3bdf7500eaf2720975d8474

Observation a999e21e-f008-46c7-bd41-fb96a7093c29 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.638677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.244479Z digest=sha256:ea5a3f1d3d5c6d6384b1ed58f71f30eb58930a5f05601b2ab6ae1f330e9e7901

Observation 64913ba6-5505-4e22-8bbf-9bb3998ba328 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.632107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.247254Z digest=sha256:fd3f7d06321387630b582eb087362f8594ce97020cc2a754f5b5fe596e8532b1

Observation 9bd36c44-6590-42cb-863f-e34b69f0c504 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.625194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.250724Z digest=sha256:5707f3e18dd4802f9f6d5cb25c938129cf9025b167c301477aebc802c4b3128a

Observation a35edc2e-255f-43c8-a8bc-2e8c9c63e1c1 · outbound

This paper cites This provides a cohesive narrative for the entire post-training pipeline, viewing it not as a sequence of disparate steps but as a unified process.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models This provides a cohesive narrative for the entire post-training pipeline, viewing it not as a sequence of disparate steps but as a unified process

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.618467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.254776Z digest=sha256:d3cce8fecd1b5f8ed2682e794ce4ef621e21b586677ea9b3ae77700951f7c31b

Observation d01692d6-8b7b-4e18-935f-1834b976261b · outbound

This paper cites winner" (preferred) response andyl is the.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models winner" (preferred) response andyl is the

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T15:49:58.330875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.258118Z digest=sha256:0d9d7f380d230263021dbbde6b021471c68b6e4e6d5e57d2f722f65b6bc94ad1

Observation 829c96dc-3465-4610-bb3e-b7248a0a1296 · outbound

This paper cites On the Diversity of Synthetic Data and its Impact on Training Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models On the Diversity of Synthetic Data and its Impact on Training Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.116911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.116911Z digest=sha256:61afce306787e87d797fdeae9a8f524b38727881df559531fd2ced41f739c06d

Pith citing papers

Observation 82e3a9ec-f72b-42da-b747-44ef9f62ab08 · inbound

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs cites this paper.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.374761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:24a92a3513b9d5f025a35f074407e4e0e7951dddffe6c08992ecf5e74aed8332