Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T22:35:35.287039Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 35 inbound Pith citation observations for arXiv:2407.16216.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T22:35:35.287039Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:39:08.244079Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
96 of 96 outbound references displayed
External citation measurements
7
pith, observed 2026-08-05T02:28:24.338817Z
Observation 41ee95c0-ed31-4565-8582-4bec1964d30a · outbound
Reinforcement Learning for LLM Post-Training: A Survey Bert: Pre-training of deep bidirectional transformers for language understanding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5ab86491-1f11-4173-8582-3f0e9a83c074 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d6f3456d-b78f-48fb-91fe-f27f3bdc1e3b · outbound
Reinforcement Learning for LLM Post-Training: A Survey Training a helpful and harmless assistant with reinforcement learning from human feedback
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8bdeb67b-9621-411b-ac3d-f21b74ef9217 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a7510842-3eba-4869-9e23-ef019229ee74 · outbound
Reinforcement Learning for LLM Post-Training: A Survey The claude 3 model family: Opus, sonnet, haiku
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 710a6eb1-3f02-4ac1-b880-ee18e4d8020b · outbound
Reinforcement Learning for LLM Post-Training: A Survey Gemini: A Family of Highly Capable Multimodal Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df9518d0-8156-496d-a784-1b0110a74a29 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Rlhf workflow: From reward modeling to online rlhf
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 19cb9fd8-1e69-4193-8b94-52e65e2b02ce · outbound
Reinforcement Learning for LLM Post-Training: A Survey Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5bedc3d-3223-4c19-9b81-8134f8cca61a · outbound
Reinforcement Learning for LLM Post-Training: A Survey Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 737657c7-ad21-433e-883a-e52a6140e94e · outbound
Reinforcement Learning for LLM Post-Training: A Survey Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5bad53b0-48e8-417a-b203-af6f6f171569 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f9169425-faf9-488b-a6c6-bb694d11a966 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Manning, and Chelsea Finn
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 90bfbf4d-0a77-4b8f-ac5e-338a80697b01 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Smaug: Fixing failure modes of preference optimisation with dpo-positive
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 12833e43-3931-4d83-814e-a1eb31bcf212 · outbound
Reinforcement Learning for LLM Post-Training: A Survey β-dpo: Direct preference optimization with dynamic β
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 74acbde2-9219-46fb-bbbd-3f62361eb642 · outbound
Reinforcement Learning for LLM Post-Training: A Survey A general theoretical paradigm to understand learning from human preferences
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ee29a6ea-2640-41c1-bd86-8919971c360d · outbound
Reinforcement Learning for LLM Post-Training: A Survey sdpo: Don’t use your data all at once
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f2cd99c0-c23d-4eae-bca1-fdd7cb9a4f4f · outbound
Reinforcement Learning for LLM Post-Training: A Survey From r to q∗: Your language model is secretly a q-function
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 37068e4e-f7d9-4b17-8516-1620c2cff80c · outbound
Reinforcement Learning for LLM Post-Training: A Survey Token-level direct preference optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e864275-507d-4ae1-9998-22abf3fcbc60 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Self-rewarding language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5ab8bc47-8da8-4d3e-99f2-450e8633bd52 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Some things are more cringe than others: Iterative preference optimization with the pairwise cringe loss
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 14e27ea6-a84f-445e-8675-73cda460eed1 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Kto: Model alignment as prospect theoretic optimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eb2c37f0-9d2b-40a5-bde0-424cc1cc69a9 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Offline regularised reinforcement learning for large language models alignment
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2dc7bb51-2f68-439e-bc2a-7eec04505a71 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Orpo: Monolithic preference optimization without reference model
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3a15b19e-901b-4da0-895f-1ff4d2d85cad · outbound
Reinforcement Learning for LLM Post-Training: A Survey Paft: A parallel training paradigm for effective llm fine-tuning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2c05bf88-20f3-4fd2-854a-c08d4b8f7cc1 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Disentangling length from quality in direct preference optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cba4f34d-36e6-4e27-8479-93be185c267e · outbound
Reinforcement Learning for LLM Post-Training: A Survey Simpo: Simple preference optimization with a reference-free reward
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d7983a5b-7419-4b44-949b-68ce7060fac5 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ff98a8ae-3b4e-4d9f-aaf1-44c676d78743 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Liu, and Xuanhui Wang
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4d7fde4-792c-4b0d-b6a0-3cddc5a58f94 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Rrhf: Rank responses to align language models with human feedback without tears
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 11a31772-b1e4-4f5d-b2c1-99ac4f709805 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Preference ranking optimization for human alignment
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e0b4a298-5c89-452c-b199-8df7ba70b1a6 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Negating negatives: Alignment without human positive samples via distributional dispreference optimization
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e143ba21-90f9-4991-871d-9f73a25da5c9 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Negative preference optimization: From catastrophic collapse to effective unlearning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aea6477c-9704-42a0-b7ed-bd91fc5afd24 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Contrastive preference optimization: Pushing the boundaries of llm performance in machine translation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 81382c5c-effc-4216-bb54-8f9c4e55463e · outbound
Reinforcement Learning for LLM Post-Training: A Survey Mankowitz, Doina Precup, and Bilal Piot
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a4f2e5ab-84a8-4f7b-8232-be7600e3cb30 · outbound
Reinforcement Learning for LLM Post-Training: A Survey A minimaximalist approach to reinforcement learning from human feedback
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d32da5e4-3c98-43c7-ba72-639be4d84958 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Direct nash optimization: Teaching language models to self-improve with general preferences
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a8b485f9-bed1-48bb-b1a1-db9f3f5814d1 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7456c3d9-499d-4871-9292-2ecdbe5a8a2b · outbound
Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a57386fa-1f0c-4055-976e-6be0bc487193 · outbound
Reinforcement Learning for LLM Post-Training: A Survey A markovian decision process
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8ab25870-c9a6-46db-82e5-00fd6a328bc9 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Hashimoto
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1778653-468c-4c57-ae85-77d0e4da3496 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Bleu: a method for automatic evaluation of machine translation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 32f7872b-193f-4e3c-b8f5-2b2cb02ff5ce · outbound
Reinforcement Learning for LLM Post-Training: A Survey Rouge: A package for automatic evaluation of summaries
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7a2547ee-ce4c-427e-bdf1-793bd3cf4150 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Weinberger, and Yoav Artzi
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 83c152fe-d6fe-4aa6-8797-04e68279cce9 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 696cf556-b69d-47cd-905b-9a70a9bd247f · outbound
Reinforcement Learning for LLM Post-Training: A Survey Truthfulqa: Measuring how models mimic human falsehoods
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4d46e208-6801-4766-a3ec-de799503b78d · outbound
Reinforcement Learning for LLM Post-Training: A Survey Chain-of-thought prompting elicits reasoning in large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 611436f5-ae08-47ce-84a4-aa10272c94d3 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b89e7f9e-2706-4d0b-b0f6-dac876fbcc1f · outbound
Reinforcement Learning for LLM Post-Training: A Survey Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 48144c6b-97a2-42f1-a558-51791ce80eea · outbound
Reinforcement Learning for LLM Post-Training: A Survey High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6e5b5ed8-83a9-4526-bca2-2fd5350c0a3b · outbound
Reinforcement Learning for LLM Post-Training: A Survey Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 985f1914-c302-4e1d-b75d-366da530af38 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Liu, and Jialu Liu
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ebebe5be-521a-4bce-bf94-a49137173aa8 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 914132a8-46b6-498b-93ee-1e777a015ab2 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Maas, Raymond E
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 44b9fd79-d89e-4ca1-88e4-5c4a81c67f8f · outbound
Reinforcement Learning for LLM Post-Training: A Survey Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c127a630-88c7-4e30-9455-117dd01e7e21 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6220a988-6323-437a-a3a2-8994ac19879a · outbound
Reinforcement Learning for LLM Post-Training: A Survey Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 93e86d94-962d-42d4-91ad-ad6bb5319b60 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Xing, Hao Zhang, Joseph E
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 45c694d0-ef68-4c86-b975-89963006acf4 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Pythia: A suite for analyzing large language models across training and scaling
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af334596-b044-46de-a42a-9ed0357a83bd · outbound
Reinforcement Learning for LLM Post-Training: A Survey Solar 10.7b: Scaling large language models with simple yet effective depth up-scaling
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7dffecc0-3606-4704-8e5d-56b776ea295d · outbound
Reinforcement Learning for LLM Post-Training: A Survey Orca: Progressive learning from complex explanation traces of gpt-4
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cbbc87f2-12e0-40f3-81a2-9b5a366cc81a · outbound
Reinforcement Learning for LLM Post-Training: A Survey Ultrafeedback: Boosting language models with high-quality feedback
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6cac6db5-fd88-47cc-8280-b276c08dc9aa · outbound
Reinforcement Learning for LLM Post-Training: A Survey Measuring massive multitask language understanding
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e9d142e5-78ac-4ab3-845a-ec23f0747172 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Winogrande: An adversarial winograd schema challenge at scale
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bf26ad90-ca31-454c-9f01-ca8cd4f79f12 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Training Verifiers to Solve Math Word Problems
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 737de695-d107-44c9-89d1-eac482791dbf · outbound
Reinforcement Learning for LLM Post-Training: A Survey Generalized preference optimization: A unified approach to offline alignment
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b3ea4c40-b1e4-47f9-9a82-ea6c2e86583e · outbound
Reinforcement Learning for LLM Post-Training: A Survey Language models are unsupervised multitask learners
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7bb92faf-99cb-49ac-a7d1-4fb093f39379 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Llama 2: Open foundation and fine-tuned chat models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 829acdd5-5433-4a4f-bc7c-8e798de4b516 · outbound
Reinforcement Learning for LLM Post-Training: A Survey The cringe loss: Learning what language not to model
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d2973944-a7d3-47e5-adff-5e0316f35255 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Hashimoto
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 07ea3b54-dd38-4bb9-ab20-65e19538d55a · outbound
Reinforcement Learning for LLM Post-Training: A Survey Advances in prospect theory: Cumulative representation of uncertainty
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 66fabb02-424f-437d-9bf7-87525f861d18 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Proximal policy optimization algorithms
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 577b2d2d-1ae5-46ed-8d3f-1497d9cb6209 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 07bff1ea-da89-472d-ad0b-613b07be387b · outbound
Reinforcement Learning for LLM Post-Training: A Survey Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1c6b788f-266b-496c-97aa-6397074e733c · outbound
Reinforcement Learning for LLM Post-Training: A Survey Phi-2: The surprising power of small language models
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f7ef9e94-119e-42f1-a00e-468bdeafab0f · outbound
Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation daeb35c3-6af9-41c4-9bfc-1681de57d7b2 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Instruction-following evaluation for large language models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a63075f8-1aa7-4c43-a0a0-a6a0848b2539 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 17579027-2b4a-43c1-bcc2-7dcb0935f694 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Gonzalez, and Ion Stoica
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3df9b7fd-5e9e-4321-a88c-f726d6c706ac · outbound
Reinforcement Learning for LLM Post-Training: A Survey Simple statistical gradient-following algorithms for connectionist reinforcement learning
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d944cac-3c8d-46d5-a075-87e31d2ec18c · outbound
Reinforcement Learning for LLM Post-Training: A Survey Buy 4 REINFORCE samples, get a baseline for free!
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2db9056c-7483-4753-8b60-19302329c864 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Learning to rank for information retrieval
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5d0e78e9-0a98-4bc1-8f8f-ccdf8b8776a1 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b00bca99-5099-4c51-b2f5-c087e9a0f940 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Openassistant conversations – democratizing large language model alignment
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation caa32368-772c-4fac-aa0e-8d37c2064eb6 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Hashimoto
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f41a71fb-eece-4c83-bb8b-db9a8deef6e5 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Secrets of rlhf in large language models part ii: Reward modeling
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 732b60fd-bad0-4dd5-b241-e2d2243d738d · outbound
Reinforcement Learning for LLM Post-Training: A Survey Safe rlhf: Safe reinforcement learning from human feedback
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a3466e86-ac7c-4a7d-a803-3cfa805308bf · outbound
Reinforcement Learning for LLM Post-Training: A Survey Lipton, and J
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9ead4ea7-08ef-4c83-b1c6-b5622e5e4d4a · outbound
Reinforcement Learning for LLM Post-Training: A Survey A paradigm shift in machine translation: Boosting translation performance of large language models
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 29ad9da9-968d-42cf-934d-f2077af82f7a · outbound
Reinforcement Learning for LLM Post-Training: A Survey On the limitations of the elo, real-world games are transitive, not additive
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 35116a69-8d42-4971-b104-f0274d074a9c · outbound
Reinforcement Learning for LLM Post-Training: A Survey Self-play preference optimization for language model alignment
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e6375735-596d-49b2-b62b-63f99fa40c54 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Schapire
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e2eb508d-fac0-49e5-9c80-046ac23e41ea · outbound
Reinforcement Learning for LLM Post-Training: A Survey Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a58d8c4c-8b7f-4c89-8fbe-e47710ed9aa2 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Orca 2: Teaching small language models how to reason
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 10e9967c-47a8-4a89-a04e-dd069ca98ca6 · outbound
Reinforcement Learning for LLM Post-Training: A Survey On decoding strategies for neural text generators
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6892be44-c36d-4e92-b3c9-77c487ee1612 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Insights into alignment: Evaluating dpo and its variants across multiple tasks
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 422cf210-88ee-448d-9e2c-acbc1f761e69 · outbound
Reinforcement Learning for LLM Post-Training: A Survey Is dpo superior to ppo for llm alignment? a comprehensive study
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b90d553c-2bc3-4b50-a060-c07269eedf3a · inbound
ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Reinforcement Learning for LLM Post-Training: A Survey
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4bb4783a-b6b6-4dad-bb60-51766d8a175a · inbound
Exploring the Secondary Risks of Large Language Models Reinforcement Learning for LLM Post-Training: A Survey
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ebd692bf-faeb-40d8-9d2e-c6ff0bde5d3a · inbound
EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention Reinforcement Learning for LLM Post-Training: A Survey
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5f5c0311-173e-4dcc-b796-221408854dff · inbound
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Reinforcement Learning for LLM Post-Training: A Survey
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fd962b78-169e-4c63-8873-44cb2333d49d · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reinforcement Learning for LLM Post-Training: A Survey
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cf9f655d-a984-48d5-a5e7-b7d55157347a · inbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Reinforcement Learning for LLM Post-Training: A Survey
Reference 273
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b7f581-ab53-4d4e-a416-17ebd7aa179b · inbound
Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial Reinforcement Learning for LLM Post-Training: A Survey
Reference 265
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a15aba6-4d84-491b-bcb5-292e9565c697 · inbound
EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models Reinforcement Learning for LLM Post-Training: A Survey
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df2adc50-b14f-4bd0-b9ea-a23274a516e8 · inbound
Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey Reinforcement Learning for LLM Post-Training: A Survey
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6300611-2eb1-4a5e-a01d-40ed2a7fe2ec · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reinforcement Learning for LLM Post-Training: A Survey
Reference 178
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b16c4e7-1b7c-4756-b793-777bda1359ce · inbound
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training Reinforcement Learning for LLM Post-Training: A Survey
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1bfb6bdd-e510-45ea-b57f-d08bcbf58d1a · inbound
Voting with the Graph: Stable RLAIF via Topological Consistency Maximization Reinforcement Learning for LLM Post-Training: A Survey
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 206d97fe-d9ff-4a62-b3a1-b31282853cbe · inbound
Maximizing the efficiency of human feedback in AI alignment: a comparative analysis Reinforcement Learning for LLM Post-Training: A Survey
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee8de5cb-22a3-44e6-a3c6-92adee7be1f8 · inbound
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing Reinforcement Learning for LLM Post-Training: A Survey
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 57d1d6d9-8bba-4db0-b197-cc43eed7a2ad · inbound
Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection Reinforcement Learning for LLM Post-Training: A Survey
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 92febfba-be07-4bce-956e-4459d2f6ea71 · inbound
VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models Reinforcement Learning for LLM Post-Training: A Survey
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a87cf072-c53d-482d-9bf8-bd7207b29052 · inbound
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training Reinforcement Learning for LLM Post-Training: A Survey
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b479ae6a-3f9f-4648-b20e-7e70fce63411 · inbound
Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization Reinforcement Learning for LLM Post-Training: A Survey
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 07619aec-be4c-422f-9420-92393b1334c8 · inbound
ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety Reinforcement Learning for LLM Post-Training: A Survey
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ad446046-50cc-4b8d-aaff-a5c5998ef8af · inbound
Pref-CTRL: Preference Driven LLM Alignment using Representation Editing Reinforcement Learning for LLM Post-Training: A Survey
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b5582d48-55c7-485b-9d8f-d0866a6e76b0 · inbound
Generating Place-Based Compromises Between Two Points of View Reinforcement Learning for LLM Post-Training: A Survey
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 70210b30-9bb9-4993-a0ae-25a3c5d1a2e2 · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models Reinforcement Learning for LLM Post-Training: A Survey
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d51e331-406f-4c13-9633-1e0b2752d58f · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models Reinforcement Learning for LLM Post-Training: A Survey
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e58f133-85e0-40f4-b5a1-a9d462a264cd · inbound
Rethinking Agentic Reinforcement Learning In Large Language Models Reinforcement Learning for LLM Post-Training: A Survey
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8cb754df-3e9a-48c1-9039-97b4dcbdc9ff · inbound
Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance Reinforcement Learning for LLM Post-Training: A Survey
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2045f3d6-b30b-407c-adac-2243f11ef719 · inbound
Agents Should Replace Narrow Predictive AI as the Orchestrator in 6G AI-RAN Reinforcement Learning for LLM Post-Training: A Survey
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ebaa50f6-de7e-4137-8a96-7d14b955712b · inbound
UNIPO: Unified Interactive Visual Explanation for RL Fine-Tuning Policy Optimization Reinforcement Learning for LLM Post-Training: A Survey
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5d940f6-bc0f-45d8-9b48-614becafa94a · inbound
Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning Reinforcement Learning for LLM Post-Training: A Survey
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 514183ef-2b4a-4335-a865-9861bc6f4aa6 · inbound
ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning Reinforcement Learning for LLM Post-Training: A Survey
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1970e252-73e9-4025-abff-429e5b4c6713 · inbound
DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment Reinforcement Learning for LLM Post-Training: A Survey
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f6851d24-0f0d-4dbb-9124-28228106628e · inbound
Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation Reinforcement Learning for LLM Post-Training: A Survey
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a7acfd67-0f89-4f61-9688-2203b6f2ddcd · inbound
Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support Reinforcement Learning for LLM Post-Training: A Survey
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8da3d815-7164-449f-9afb-35f55a83bc06 · inbound
Meta-Learning Preferences for Multilingual LLM Alignment Reinforcement Learning for LLM Post-Training: A Survey
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59a1edaa-924c-4831-a5be-ff2c11c0bace · inbound
Sound Probabilistic Safety Bounds for Large Language Models Reinforcement Learning for LLM Post-Training: A Survey
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6cadd64-ae02-4da7-8972-71698e65181d · inbound
On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment Reinforcement Learning for LLM Post-Training: A Survey
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.