Pith. sign in

Paper Citation Record · LEDGER

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models

As of 16 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2509.25148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.25148 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:49:58.258118Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T00:48:56.892634Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T21:18:58.373323Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved47
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f85890b8-4fab-4908-b811-2d4220c38e01 · outbound

This paper cites Harnessing the power of llms in practice: A survey on chatgpt and beyond.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Harnessing the power of llms in practice: A survey on chatgpt and beyond

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.792700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.061817Z digest=sha256:77ca58230709a74b4a3e00a2aad3cdc43407e6e48f81a7fe92978cfe26e884e8

Observation 7e4e84ab-65ee-43a4-9e33-dee86a050071 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.065373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.065373Z digest=sha256:292c682c92b42d7faca68a19dd95b5495162744d4ed8e1c179122fad2b12c259

Observation d9048366-9d1e-4557-b7db-57d92e51a4dc · outbound

This paper cites \texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models \texttt{R$^\textbf{2}$AI}: Towards Resistant and Resilient AI in an Evolving World

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T15:49:58.605246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.068948Z digest=sha256:ca82220ac1721fa223ca3ea3309aaa6eecd1809ba319608fdf5c8b4dc2b6643f

Observation 3e67ee99-bf16-4aeb-b499-937ff5aa3806 · outbound

This paper cites A survey on post-training of large language models.arXiv e-prints, pages arXiv–2503, 2025.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models A survey on post-training of large language models.arXiv e-prints, pages arXiv–2503, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.785712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.072181Z digest=sha256:07ac8db218e031ac328a3516fbd549b11fed23d334e3d23384e6f8d7e9d033cd

Observation 4fdda5c9-fb6f-4973-82ea-31fa4c32784f · outbound

This paper cites Training language models to follow instructions with human feedback.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Training language models to follow instructions with human feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.075991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.075991Z digest=sha256:1bbd74f609cd0914a7236d743a702feb50e357c6a3e261a75c5513825a26aeff

Observation 4f8a0d64-a97c-4051-83e3-f8f4f95bf819 · outbound

This paper cites Deep reinforcement learning from human preferences.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Deep reinforcement learning from human preferences

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.078529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.078529Z digest=sha256:927d023ddd2bbca5bc557377b77df81ea24dce4384e63f6aeebf2880d06a0e57

Observation d98b61c7-d0b4-4444-8999-b0f0948a7d50 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.081745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.081745Z digest=sha256:9904f1cdda91e2876a22f4644397c586c93c207a46709f3655549af522c24035

Observation c8aec489-bfc7-4e0d-9795-013e3887e282 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.084732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.084732Z digest=sha256:3b2fd1af28784bd59ba7973f6193d5dcaee7ad1434a831e6fe472cf46412d233

Observation 56686d20-3a5d-490d-bb73-6df25066f822 · outbound

This paper cites Curriculum offline imitating learning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Curriculum offline imitating learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.768174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.088393Z digest=sha256:c9e628c9b2c6518e1e68dfa96994443020355214da0f8f95e38057c977ce1c47

Observation 11144b59-b2cf-4452-a379-532cd0493a36 · outbound

This paper cites Offline imitation learning with suboptimal demonstrations via relaxed distribution matching.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Offline imitation learning with suboptimal demonstrations via relaxed distribution matching

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.761477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.091500Z digest=sha256:fce5efed65c5d78edac2d7afb3fb7311139e37e3a1d5d16c15134a6bba0727d4

Observation e38fa1cb-b58c-4008-acaa-fc7a3123f3b3 · outbound

This paper cites Offline Reinforcement Learning for LLM Multi-Step Reasoning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.094264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.094264Z digest=sha256:2289c0f041f4b9c4a97bb7ba1cb96a2c1e40686215191a2123f266358c739040

Observation 1cd45998-58e6-4d1a-b03b-681575012957 · outbound

This paper cites Grounding large language models in interactive environments with online reinforcement learning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Grounding large language models in interactive environments with online reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.097622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.097622Z digest=sha256:11a76567aaef96c1a826e0beb4fc60f24c5e8a614f70194f62c46972e36a792c

Observation a34bb80f-a6e5-47ef-b26c-5de71dbef4dc · outbound

This paper cites Provably mitigating overoptimization in rlhf: Your sft loss is implicitly an adversarial regularizer.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Provably mitigating overoptimization in rlhf: Your sft loss is implicitly an adversarial regularizer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.751138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.100869Z digest=sha256:ada62b8da8bfa6f71a1a0b623351e33dd2f950ba8f670a061508180cda8341d4

Observation 12dcd533-12d3-4efd-8ae2-85c5e8b1fe88 · outbound

This paper cites Preserving Diversity in Supervised Fine-Tuning of Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Preserving Diversity in Supervised Fine-Tuning of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.103306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.103306Z digest=sha256:e1ecd99c32adfaf9115bd7b95d78f6fe63f2e639cfc7984212be602d4343167a

Observation 070ec0c1-259d-4d4b-8a41-30d2908456ab · outbound

This paper cites Regularizing Neural Networks by Penalizing Confident Output Distributions.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Regularizing Neural Networks by Penalizing Confident Output Distributions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.106522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.106522Z digest=sha256:525ea2ca800d389c4fb42416b4ad370e7a0b9ec3527354bdee4c88b79b2d5367

Observation b801e931-70b3-47d3-9822-e7b9710433c2 · outbound

This paper cites Stanford alpaca: An instruction-following llama model, 2023.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Stanford alpaca: An instruction-following llama model, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.743340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.109290Z digest=sha256:29c5ac2f150e16dd1755bcd673157e75229e4d3199ab96349b443e4a4d2c7399

Observation 2b49f18d-da45-43a6-b2be-234497283dd4 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.111627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.111627Z digest=sha256:6e408cc46d765bdbf8005b10dfe9c63dd4b6162902a4ad6b023064b3f95d557e

Observation 2316bf61-624d-421a-83bc-f0cf68cc97f7 · outbound

This paper cites On the diversity of synthetic data and its impact on training large language models,.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models On the diversity of synthetic data and its impact on training large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.736182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.114383Z digest=sha256:24d21c6b394b1ab491ce773a5be0c93b0ebe41a586eb4de295dc0a701f00c8bd

Observation b8b1e8c8-d5b7-413e-9e0c-7d90f7b60170 · outbound

This paper cites Condor: Enhance llm alignment with knowledge-driven data synthesis and refinement.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Condor: Enhance llm alignment with knowledge-driven data synthesis and refinement

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.728110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.119357Z digest=sha256:97cacce958286c6dbdf02b9a8822f2e835cd7402f5686197017e3104f966ba73

Observation 938e0ef4-961e-4a9c-9218-bd996afa1a9b · outbound

This paper cites Idgen: Item discrimination induced prompt generation for llm evaluation.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Idgen: Item discrimination induced prompt generation for llm evaluation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.721234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.123053Z digest=sha256:6c8aabb988ce5bc6f0a177dac3f219ed2fad3cc150ed0eb5cac121fdea039d15

Observation d27af825-ab83-4627-be9a-6228d708508d · outbound

This paper cites Qwen3 Technical Report.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Qwen3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.126327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.126327Z digest=sha256:1d8ab0862a0022eb43c155c88cda53917d357552be0d5c2599b2ef63bdaa1053

Observation 08e3b7df-4394-48c2-85f5-199226c1483e · outbound

This paper cites DeepSeek-V3 Technical Report.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DeepSeek-V3 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.129064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.129064Z digest=sha256:eacda120b93a234f6d89eb4e0da6cff0a649a10fd9e9f34ec71b4dca8c9a87d9

Observation 4fcddce0-5c59-4725-bca2-b15566824e6e · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Gemma: Open Models Based on Gemini Research and Technology

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.132125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.132125Z digest=sha256:96850bf204eebf90bce1f32a6be5c1d27ecb41dfeb06a5627e614720769ef52d

Observation 98c886ce-718d-4454-8c6e-ccafe91ea064 · outbound

This paper cites GPT-4 Technical Report.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.136271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.136271Z digest=sha256:79e0b7aefe40bbbd5782ec9f3a1438f7ac327b59719633615b4f34fff446d8ac

Observation 23866afd-1047-4d4b-b9ec-a3a32e138b27 · outbound

This paper cites Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.138798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.138798Z digest=sha256:b13fe98aa9a07d6535e346ad499a18b474c82a5add1dbee9c201914e2eefe0c0

Observation d439eeeb-0b36-4daa-8288-a0f1779a44bd · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.142255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.142255Z digest=sha256:98ce43fa785197c7b105d20b2d050e6570de7eb15b16bdf490c42b51db3c5744

Observation df333a4f-9aea-4408-801c-1d20fc8e213d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.146538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.146538Z digest=sha256:648918a725e93774e059dc60d2873967756d236eea581299a0d71b40997fb451

Observation 7c7ec438-faeb-4e0b-ba3b-1a95d3c8c691 · outbound

This paper cites Reinforcement learning with verifiable rewards: Grpo’s effective loss, dynam- ics, and success amplification.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Reinforcement learning with verifiable rewards: Grpo’s effective loss, dynam- ics, and success amplification

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.150276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.150276Z digest=sha256:a8e933ee96d6df8e8083ef171562bfcc323756f0b19e7194f7f60dba555c3776

Observation 028c0ad8-5942-4d92-9c1f-aaa47d435aab · outbound

This paper cites AutoGLM: Autonomous Foundation Agents for GUIs.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models AutoGLM: Autonomous Foundation Agents for GUIs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.152922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.152922Z digest=sha256:4722f94fd619b800c56018508a324494ebe9d7e4a34b88daa9f23354562f907b

Observation 7ebdcc37-7fe3-49c2-b0a6-2d60a82ad4c6 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.155753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.155753Z digest=sha256:05bd12c10e543f3dbc2cc2cff98305bbd9691e540582da67daf8d407b13c003f

Observation a7faba00-97d3-403c-a438-df0d71050587 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.158558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.158558Z digest=sha256:594b14bd26b707feacf9a823c39e5d70d7089793e96df466007aef1dd2d06516

Observation 04c91798-d270-49d5-8e48-98e58360054b · outbound

This paper cites Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.162168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.162168Z digest=sha256:5f66b3999e1eaa2bc93d83018b49e9529a534abc7f363247f1a88645f0027c5f

Observation 98ef0e3a-b85d-4bb8-9e5d-8a6ec1c0af87 · outbound

This paper cites On-policy rl meets off-policy experts: Harmonizing supervised fine- tuning and reinforcement learning via dynamic weighting.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models On-policy rl meets off-policy experts: Harmonizing supervised fine- tuning and reinforcement learning via dynamic weighting

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.165261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.165261Z digest=sha256:ca4eb999072c46141ea1bac3d0654acc37a79f8b3fd4eedd9a7ff9a1f0cf0668

Observation b3ecce42-6552-4fbe-b36c-1906c1d830e9 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Learning to Reason under Off-Policy Guidance

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.169948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.169948Z digest=sha256:7048261c9d4767170cddde205bcccdaa52fb86313c23d384cddf7a920ee5dcc1

Observation 8858a5b0-372f-42e1-a199-4c9d910cbcb6 · outbound

This paper cites BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.173572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.173572Z digest=sha256:ca407f27ef5a53f283bd314b6c1efab5f2f0786aabf624926c706af3e4d7ad97

Observation a5c519e1-9d20-4eb9-af14-6b057420aa87 · outbound

This paper cites SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.176602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.176602Z digest=sha256:20341a8f0c54d566f6412842ef023bdef54e186b003651768c94536cc688c38e

Observation 875a2f1d-728a-487a-90bc-606c0c1d27da · outbound

This paper cites Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.179658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.179658Z digest=sha256:177033f03875d82f66294e90c68c879464023448d68ecb636dc1c115cb748832

Observation a198c5bc-aa3d-440b-bb9a-8d46304380f5 · outbound

This paper cites Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing, 2024.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.183190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.183190Z digest=sha256:9c1c19a213f54c97b14a96cc68aef2da87c1288c92eaca3359e8ebf9f4684a57

Observation 54df15d6-af5d-4049-993c-1bde54bbe74c · outbound

This paper cites Generalizing verifiable instruction following, 2025.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Generalizing verifiable instruction following, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.185865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.185865Z digest=sha256:0bf4693e474187eb958d3ff2e7d4d7d0654382850cec8f60c21873fc35b0d4b8

Observation a762c342-dd82-4265-8ff4-5eedfda017f5 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Instruction-Following Evaluation for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.188396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.188396Z digest=sha256:0db450e3db7009f1124192d178abceac04f866adf19252e264055fe6b04943cd

Observation 92da0776-98c2-4f65-8985-74d5e47917c4 · outbound

This paper cites Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.190607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.190607Z digest=sha256:08bbadf49f8cc0a46f1b21fe5c3b076b9cc219303ff1d1bfb98b406281b0b070

Observation f85c6f8f-bac3-47e1-b658-f4e938b68f5b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Measuring Massive Multitask Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.193411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.193411Z digest=sha256:48b0255bdf586a7bf055f5c07782fec08d150b79f2956733e7a25ac3c7d1222c

Observation 9ee6b06f-52d5-4c6c-b909-a4c1087c866c · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.196925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.196925Z digest=sha256:ddee603f1f16046d8ef7974804f5dd7da627feef8e842770e93c91ae3b7594b9

Observation 804d1d4c-4045-4555-a32d-91e48d1258b3 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Gpqa: A graduate-level google-proof q&a benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.199271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.199271Z digest=sha256:51b9176f20324253ed49e0a1d934645660c0d57af2d654e4eac9e6965638d5f9

Observation b947f471-b0f1-400f-af49-1deb956d0934 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Evaluating Large Language Models Trained on Code

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.201429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.201429Z digest=sha256:4a5235a6a63bebd669c8b7c2bff90640689bc1f934efccd00a16b21f9eb5b506

Observation 0979e746-9fab-4609-beee-71afeb3e1c52 · outbound

This paper cites Program Synthesis with Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Program Synthesis with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.205422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.205422Z digest=sha256:b9cb46c55447bdcfdba6b852a206c3f609207515c099157f9fa9e0a5d8496b95

Observation 6005055c-4816-4fd3-97ff-ab9624d64ae7 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.209138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.209138Z digest=sha256:b6448a64ba84fcefe243aa2774fa353954df85e70ba2ea9218f5956667d92603

Observation b7ace46a-de22-4611-9396-116de9b0e6f1 · outbound

This paper cites Let’s verify step by step.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Let’s verify step by step

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.212019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.212019Z digest=sha256:fd9638f438ccdead2b86685d20c381b4534d4312406dc9fe18ec29c8b721c50b

Observation 82d5c9da-f8f5-4afa-8ba3-be80aee45ae7 · outbound

This paper cites Theoremqa: A theorem-driven question answering dataset.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Theoremqa: A theorem-driven question answering dataset

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.699052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.214895Z digest=sha256:0e5d9ec4782c8a9432565343b926e67d361f32a45110a101dfb3522ed7b7edd8

Observation 62e6de16-ca27-4c2d-9a22-65beb8e7d1d8 · outbound

This paper cites Cmmlu: Measuring massive multitask language understanding in chinese, 2023.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Cmmlu: Measuring massive multitask language understanding in chinese, 2023

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.217412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.217412Z digest=sha256:2f28a65fb5f58855b50eb3d9bef857c0b666f5f44ee86c7042aa675a5a852ac0

Observation e8fcaa28-bd44-49ef-b07e-4760195a12b7 · outbound

This paper cites C-eval: 12 A multi-level multi-discipline chinese evaluation suite for foundation models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models C-eval: 12 A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.688924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.219886Z digest=sha256:f69cf09589c81a44682b4c885f1e3a03cf33151160c8dc600f7cd14635b8b618

Observation ebb3d37d-99fb-4248-8f04-9148819eb483 · outbound

This paper cites Pre-trained policy discriminators are general reward models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Pre-trained policy discriminators are general reward models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.222765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.222765Z digest=sha256:c6fb3a38ef6baf8da87a2c86e2591397faaf144e1e2e9b9d104e82dc5a22d2df

Observation 324a8893-47d2-4c6a-a407-5dbb8abfa522 · outbound

This paper cites Swift:a scalable lightweight infrastructure for fine-tuning, 2025.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Swift:a scalable lightweight infrastructure for fine-tuning, 2025

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.681698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.225967Z digest=sha256:53ebe1af2d001b399140ac6717edb4608bf9b668ba23f8c2fed01a90ea47376a

Observation ac153536-d7aa-48c5-9f68-870b934ff8ad · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Hybridflow: A flexible and efficient rlhf framework

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.228336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.228336Z digest=sha256:450b28a7b82c03e1128b788f22e60d26f8ea182761c1d4d94085fac63d09f206

Observation f6522c83-15cd-4dc4-8a16-7f7d405c00b9 · outbound

This paper cites preference.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models preference

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.671519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.230938Z digest=sha256:b4198a9b38b64fcbd8558c274a97307dcc4383da467c67a065cda3391b9fc62e

Observation 5c6a21b2-369d-4393-95a4-8a8e50d8b392 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.663350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.234960Z digest=sha256:2e0bd56cb5d8ea8f20ba65dcce15cd49f42aa313aefb2f947fbbe4cc11f267d0

Observation b4ef34fe-e198-4a5d-950e-33621b0407d7 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.655942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.238143Z digest=sha256:5f58868dfa7bce57d2ee7f35db0ff4a5e38dd2916efff187fe1cace9db6572de

Observation 35491afa-7d8e-49b3-b244-848509f433bb · outbound

This paper cites Due to the properties of the Dirac delta function, the integral is non-zero only at the single point y=y ∗.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Due to the properties of the Dirac delta function, the integral is non-zero only at the single point y=y ∗

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.646748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.240961Z digest=sha256:536b254c1ddb7c96a2c559bbe3c730d597663031a06892eb7af35b28974ad90e

Observation a999e21e-f008-46c7-bd41-fb96a7093c29 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.638677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.244479Z digest=sha256:3213d88d6a45dbb5d6d720bd23ee7dcbfb020040058f84ca53d6145342d486d3

Observation 64913ba6-5505-4e22-8bbf-9bb3998ba328 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.632107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.247254Z digest=sha256:8a573ccb425d38dcd89175013e2cc6943c3e0f267cc2570ca0a48f51895c05bb

Observation 9bd36c44-6590-42cb-863f-e34b69f0c504 · outbound

This paper cites an unresolved cited work.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:49:58.625194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.250724Z digest=sha256:abad7c8957da9bec6bd5c0da79ce9e3b0b185e0c37c90b5bebdf6fd4402017cb

Observation a35edc2e-255f-43c8-a8bc-2e8c9c63e1c1 · outbound

This paper cites This provides a cohesive narrative for the entire post-training pipeline, viewing it not as a sequence of disparate steps but as a unified process.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models This provides a cohesive narrative for the entire post-training pipeline, viewing it not as a sequence of disparate steps but as a unified process

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:49:58.618467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.254776Z digest=sha256:d8b9c766f7e782f45036b069601f5dc8573ce07b15cec267fce6db89d6e6ee00

Observation d01692d6-8b7b-4e18-935f-1834b976261b · outbound

This paper cites winner" (preferred) response andyl is the.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models winner" (preferred) response andyl is the

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T15:49:58.330875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:49:58.258118Z digest=sha256:c9d4704b8aeab2b07c5a3c411d9c2094ea094abed51920d5379ce74e5b6c264a

Observation 829c96dc-3465-4610-bb3e-b7248a0a1296 · outbound

This paper cites On the Diversity of Synthetic Data and its Impact on Training Large Language Models.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models On the Diversity of Synthetic Data and its Impact on Training Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.116911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.116911Z digest=sha256:9d43f3c975d66792670de32f63d513da7cb41208bc5815fd9707eb8b9335560a

Pith citing papers

Observation 82e3a9ec-f72b-42da-b747-44ef9f62ab08 · inbound

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs cites this paper.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.374761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:55586fafdc0133346666f2b2bdfeddf018b82562b83ccce278ab95049399b833