Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning for LLM Post-Training: A Survey

As of 22 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 58 inbound Pith citation observations for arXiv:2407.16216.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.16216 v4

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T22:35:35.287039Z

measured 154 of 154 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 58 of 58 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:12.669701Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

96 of 96 outbound references displayed

  • verified exact6
  • verified fuzzy81
  • unresolved8
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

7
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 41ee95c0-ed31-4565-8582-4bec1964d30a · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Reinforcement Learning for LLM Post-Training: A Survey Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.479044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:2118fbc222a0a604d79b9ef35dafd576a7df132a132c05355744f4576dc5c8a5

Observation 5ab86491-1f11-4173-8582-3f0e9a83c074 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.482409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:04aa84b11dc03652822b882eb44056273fbcdd68d4c5b772be4b2ed15d4ed2ce

Observation d6f3456d-b78f-48fb-91fe-f27f3bdc1e3b · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback.

Reinforcement Learning for LLM Post-Training: A Survey Training a helpful and harmless assistant with reinforcement learning from human feedback

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.486346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:a450d8692c8d400339542d82c4db151f5df9c0895539aa1b7203f51f99e2f59b

Observation 8bdeb67b-9621-411b-ac3d-f21b74ef9217 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.416492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:c18c6b59831d5eb478860c2dfc2401aa79385e653cb4268c6c7a8dc48ac22f36

Observation a7510842-3eba-4869-9e23-ef019229ee74 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Reinforcement Learning for LLM Post-Training: A Survey The claude 3 model family: Opus, sonnet, haiku

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.475329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:843cd271454cd4f0cd9ba3123ef66631da478ca194cd76ed9c1162a46f8c0fb6

Observation 710a6eb1-3f02-4ac1-b880-ee18e4d8020b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Reinforcement Learning for LLM Post-Training: A Survey Gemini: A Family of Highly Capable Multimodal Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:35:50.304630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:8d9742a8fcca5aa1f27080bc8639134b17a8e02139ebca56c56f14ecdfed731e

Observation df9518d0-8156-496d-a784-1b0110a74a29 · outbound

This paper cites Rlhf workflow: From reward modeling to online rlhf.

Reinforcement Learning for LLM Post-Training: A Survey Rlhf workflow: From reward modeling to online rlhf

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.378072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:62cd46a6fc63cf1e7fa524b46b97d2f4b90b5c68efc9ef7ff1aff8b11a391e7a

Observation 19cb9fd8-1e69-4193-8b94-52e65e2b02ce · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint.

Reinforcement Learning for LLM Post-Training: A Survey Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.397901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:32370cab0931018b92240c1f533554c510e1b62396b893d6fc7cd361a74bd65d

Observation a5bedc3d-3223-4c19-9b81-8134f8cca61a · outbound

This paper cites Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan.

Reinforcement Learning for LLM Post-Training: A Survey Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.639341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:84843a35e5c593e425d792bc02e1523caaaed8496be2bd2d92e488fef76c6434

Observation 737657c7-ad21-433e-883a-e52a6140e94e · outbound

This paper cites Rlaif: Scaling reinforcement learning from human feedback with ai feedback.

Reinforcement Learning for LLM Post-Training: A Survey Rlaif: Scaling reinforcement learning from human feedback with ai feedback

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.573067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:f47cc308e620feff7713c149f884ebcc2cc4e56bafcdca7912dac785847c58f8

Observation 5bad53b0-48e8-417a-b203-af6f6f171569 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.583731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:05ea66d28e6ed95fdfcaa38fdf987b944197a0935e949068150f9823c5a2d973

Observation f9169425-faf9-488b-a6c6-bb694d11a966 · outbound

This paper cites Manning, and Chelsea Finn.

Reinforcement Learning for LLM Post-Training: A Survey Manning, and Chelsea Finn

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.545251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:cfda469841dc3e1bcd8f8c0612f6b9729eb6ecca9852c399f714da899334994c

Observation 90bfbf4d-0a77-4b8f-ac5e-338a80697b01 · outbound

This paper cites Smaug: Fixing failure modes of preference optimisation with dpo-positive.

Reinforcement Learning for LLM Post-Training: A Survey Smaug: Fixing failure modes of preference optimisation with dpo-positive

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.539904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:920d57ef83a75db538b9a519b60696cbdbf61bec4142d448442bfddbe353906d

Observation 12833e43-3931-4d83-814e-a1eb31bcf212 · outbound

This paper cites β-dpo: Direct preference optimization with dynamic β.

Reinforcement Learning for LLM Post-Training: A Survey β-dpo: Direct preference optimization with dynamic β

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.549126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:cf8400cfaa6ec246b19c8b721d6929406888d19b107c5ffdac234635238939c4

Observation 74acbde2-9219-46fb-bbbd-3f62361eb642 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Reinforcement Learning for LLM Post-Training: A Survey A general theoretical paradigm to understand learning from human preferences

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.565905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:b17c9c913944291b89b0af2fa5e20cf7da2ef98961acd88a2e89ce225527bedc

Observation ee29a6ea-2640-41c1-bd86-8919971c360d · outbound

This paper cites sdpo: Don’t use your data all at once.

Reinforcement Learning for LLM Post-Training: A Survey sdpo: Don’t use your data all at once

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.615981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:9730e88a6d6a5205785b2fdd8fa5b857e2fa25d08633ec0b4cffbfb71c0e5584

Observation f2cd99c0-c23d-4eae-bca1-fdd7cb9a4f4f · outbound

This paper cites From r to q∗: Your language model is secretly a q-function.

Reinforcement Learning for LLM Post-Training: A Survey From r to q∗: Your language model is secretly a q-function

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.693376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:3376ce00c797af64d8915bd438c1435df8438a5202e78ee319c2f6b48db02b11

Observation 37068e4e-f7d9-4b17-8516-1620c2cff80c · outbound

This paper cites Token-level direct preference optimization.

Reinforcement Learning for LLM Post-Training: A Survey Token-level direct preference optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.493725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:bd36de529099f57dcb41df5389f7c4f1f7a68e6161860947ab426123b1f439d7

Observation 2e864275-507d-4ae1-9998-22abf3fcbc60 · outbound

This paper cites Self-rewarding language models.

Reinforcement Learning for LLM Post-Training: A Survey Self-rewarding language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.470156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:0e28ade4fe82b015733697906f1db2d3e58a06e0068b059872249c2a2909cc8e

Observation 5ab8bc47-8da8-4d3e-99f2-450e8633bd52 · outbound

This paper cites Some things are more cringe than others: Iterative preference optimization with the pairwise cringe loss.

Reinforcement Learning for LLM Post-Training: A Survey Some things are more cringe than others: Iterative preference optimization with the pairwise cringe loss

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.498202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:074d9620e30e7f2f2da03d11d2a6d673bdf1aa4ae7dd6a21241ab4dd2bc23f0a

Observation 14e27ea6-a84f-445e-8675-73cda460eed1 · outbound

This paper cites Kto: Model alignment as prospect theoretic optimization.

Reinforcement Learning for LLM Post-Training: A Survey Kto: Model alignment as prospect theoretic optimization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.531924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:207cebcf6208b8baa8b3b8246ca08d6d6ebe665af423b3dba8f2a016cbe646ef

Observation eb2c37f0-9d2b-40a5-bde0-424cc1cc69a9 · outbound

This paper cites Offline regularised reinforcement learning for large language models alignment.

Reinforcement Learning for LLM Post-Training: A Survey Offline regularised reinforcement learning for large language models alignment

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.514400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:ec8cb17d90b3938b0d98c49195077db9a325cca97565e6a6cf5539c9c3516f6a

Observation 2dc7bb51-2f68-439e-bc2a-7eec04505a71 · outbound

This paper cites Orpo: Monolithic preference optimization without reference model.

Reinforcement Learning for LLM Post-Training: A Survey Orpo: Monolithic preference optimization without reference model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.507139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:c04e21f710057d527df86f8ad89248a647e05971a92de86fdaca97c75602a0ff

Observation 3a15b19e-901b-4da0-895f-1ff4d2d85cad · outbound

This paper cites Paft: A parallel training paradigm for effective llm fine-tuning.

Reinforcement Learning for LLM Post-Training: A Survey Paft: A parallel training paradigm for effective llm fine-tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.462277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:adda4182c00af80433a90124a33516ef38525846f2d7b64e93140713c449f59d

Observation 2c05bf88-20f3-4fd2-854a-c08d4b8f7cc1 · outbound

This paper cites Disentangling length from quality in direct preference optimization.

Reinforcement Learning for LLM Post-Training: A Survey Disentangling length from quality in direct preference optimization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.489672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:c526409d82634cbef09a618a8245a10fc682710a9e90f589a65392b855fe2f58

Observation cba4f34d-36e6-4e27-8479-93be185c267e · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

Reinforcement Learning for LLM Post-Training: A Survey Simpo: Simple preference optimization with a reference-free reward

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.524829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:fb7376a86070141a24f665acd2158c5fc596afc3eb75aec596d50128f91820c2

Observation d7983a5b-7419-4b44-949b-68ce7060fac5 · outbound

This paper cites Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms.

Reinforcement Learning for LLM Post-Training: A Survey Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.441352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:a11bdf1ebea8c45d1a93b42efc8ed0a279d105eb797561b8db2cef1cca064397

Observation ff98a8ae-3b4e-4d9f-aaf1-44c676d78743 · outbound

This paper cites Liu, and Xuanhui Wang.

Reinforcement Learning for LLM Post-Training: A Survey Liu, and Xuanhui Wang

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.559480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:d3ad26fdf240f66ec654a8ab3968675646d732d28f2d4092289301be9fcd9531

Observation d4d7fde4-792c-4b0d-b6a0-3cddc5a58f94 · outbound

This paper cites Rrhf: Rank responses to align language models with human feedback without tears.

Reinforcement Learning for LLM Post-Training: A Survey Rrhf: Rank responses to align language models with human feedback without tears

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.580163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:813ce2623db485d567b107e4ab3873e3269ad2249cfc0ecbf7ade772a1b4151c

Observation 11a31772-b1e4-4f5d-b2c1-99ac4f709805 · outbound

This paper cites Preference ranking optimization for human alignment.

Reinforcement Learning for LLM Post-Training: A Survey Preference ranking optimization for human alignment

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.503370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:367217fa254e5b660d6dbe2e69746cd86b0274cbc82cb4dfb063287118105aed

Observation e0b4a298-5c89-452c-b199-8df7ba70b1a6 · outbound

This paper cites Negating negatives: Alignment without human positive samples via distributional dispreference optimization.

Reinforcement Learning for LLM Post-Training: A Survey Negating negatives: Alignment without human positive samples via distributional dispreference optimization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.419438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:b6f68e267ca635b2668e6e95594ae9366a5959565133bbf443d876c0fd4bf819

Observation e143ba21-90f9-4991-871d-9f73a25da5c9 · outbound

This paper cites Negative preference optimization: From catastrophic collapse to effective unlearning.

Reinforcement Learning for LLM Post-Training: A Survey Negative preference optimization: From catastrophic collapse to effective unlearning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.511504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:22eeb8379d2cf36aa57f558fe3e2a725e109be4961480de2ac7ac6cfe9039b2f

Observation aea6477c-9704-42a0-b7ed-bd91fc5afd24 · outbound

This paper cites Contrastive preference optimization: Pushing the boundaries of llm performance in machine translation.

Reinforcement Learning for LLM Post-Training: A Survey Contrastive preference optimization: Pushing the boundaries of llm performance in machine translation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.536251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:574a1a2dd471fce9683e341b6c99af9b2582b0fd5ec8548070e6bef50a6d4799

Observation 81382c5c-effc-4216-bb54-8f9c4e55463e · outbound

This paper cites Mankowitz, Doina Precup, and Bilal Piot.

Reinforcement Learning for LLM Post-Training: A Survey Mankowitz, Doina Precup, and Bilal Piot

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.510618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:bc376630121e1d50359f694b9251cdcce20bb8a32a30bb586f361780af60820c

Observation a4f2e5ab-84a8-4f7b-8232-be7600e3cb30 · outbound

This paper cites A minimaximalist approach to reinforcement learning from human feedback.

Reinforcement Learning for LLM Post-Training: A Survey A minimaximalist approach to reinforcement learning from human feedback

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.415463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:45669b2a370ee104d3a6ffeac935ec806bd840e33bae7dbd25f2c6f2f05e563d

Observation d32da5e4-3c98-43c7-ba72-639be4d84958 · outbound

This paper cites Direct nash optimization: Teaching language models to self-improve with general preferences.

Reinforcement Learning for LLM Post-Training: A Survey Direct nash optimization: Teaching language models to self-improve with general preferences

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.401392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:e3a04cce080e94ca477a1dafa0e8308c9328dc9b0fb11162769f96a4303ecd40

Observation a8b485f9-bed1-48bb-b1a1-db9f3f5814d1 · outbound

This paper cites Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints.

Reinforcement Learning for LLM Post-Training: A Survey Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.408714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:cd7291110996cbb1b71013122203eafcb3789e99a7b75f3f2c80429bd4b95787

Observation 7456c3d9-499d-4871-9292-2ecdbe5a8a2b · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.391681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:e1cdd5f8b94349a230333cb0e22f0c9393b4d302196188c0cf32579b9ffa84f0

Observation a57386fa-1f0c-4055-976e-6be0bc487193 · outbound

This paper cites A markovian decision process.

Reinforcement Learning for LLM Post-Training: A Survey A markovian decision process

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.388440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:0b1addf7b74f153a7142b6e100e06cff1491daed151ab298ac6294d4830008a1

Observation 8ab25870-c9a6-46db-82e5-00fd6a328bc9 · outbound

This paper cites Hashimoto.

Reinforcement Learning for LLM Post-Training: A Survey Hashimoto

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.394727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:ba4520561f6fd04e52341705ec5c23c955c25cc18cc458b34e7e2f6c8e8443ea

Observation a1778653-468c-4c57-ae85-77d0e4da3496 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Reinforcement Learning for LLM Post-Training: A Survey Bleu: a method for automatic evaluation of machine translation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.412117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:a9808a2eef6eec7ea8e60feea254218eb9c947a1c3ed044920c83c7e74bbacc6

Observation 32f7872b-193f-4e3c-b8f5-2b2cb02ff5ce · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Reinforcement Learning for LLM Post-Training: A Survey Rouge: A package for automatic evaluation of summaries

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.482168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:c8be341b9b386076eeeffc4ed7802840c98b772785bf4eca7a543ca6b42a84af

Observation 7a2547ee-ce4c-427e-bdf1-793bd3cf4150 · outbound

This paper cites Weinberger, and Yoav Artzi.

Reinforcement Learning for LLM Post-Training: A Survey Weinberger, and Yoav Artzi

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.661237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:6a00370eabdddec46c2d470afb0920382b62bb29a51cefb64b99b90870b1efc5

Observation 83c152fe-d6fe-4aa6-8797-04e68279cce9 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.429877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:7c7814d401d0b8ddf1a493a0e8491a816519542fd35d9a8e74ad0b4f7230dd30

Observation 696cf556-b69d-47cd-905b-9a70a9bd247f · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

Reinforcement Learning for LLM Post-Training: A Survey Truthfulqa: Measuring how models mimic human falsehoods

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.442972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:406725bcee7674dcada990dd1d3d7aece37ee6facdffcd7796990804d72be4d0

Observation 4d46e208-6801-4766-a3ec-de799503b78d · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Reinforcement Learning for LLM Post-Training: A Survey Chain-of-thought prompting elicits reasoning in large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.446993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:4dcaf815f9ede99123fbb71539abd6421afdaed3005dffdbed491efdc26aef9e

Observation 611436f5-ae08-47ce-84a4-aa10272c94d3 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano.

Reinforcement Learning for LLM Post-Training: A Survey Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.665586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:40d7305f8922703bf055cc057ed88966f756b9fe2026f92a6a83eb29f194be6e

Observation b89e7f9e-2706-4d0b-b0f6-dac876fbcc1f · outbound

This paper cites Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H.

Reinforcement Learning for LLM Post-Training: A Survey Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.576240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:d7c168b3a9b7365f9b9c39f5aeccff2f52925f3a2b6cb01e91b84e7f6f8b4da7

Observation 48144c6b-97a2-42f1-a558-51791ce80eea · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Reinforcement Learning for LLM Post-Training: A Survey High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:35:50.309976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:e4f22efab22fcbcde06fb03404ea022300e5c7f9dc5572a008dd28f6e84cee22

Observation 6e5b5ed8-83a9-4526-bca2-2fd5350c0a3b · outbound

This paper cites Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J.

Reinforcement Learning for LLM Post-Training: A Survey Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.444955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:9b25651c3f3989432bd09ee8968bcfe4fb1cf1a990b9905474af1b97e668e5c4

Observation 985f1914-c302-4e1d-b75d-366da530af38 · outbound

This paper cites Liu, and Jialu Liu.

Reinforcement Learning for LLM Post-Training: A Survey Liu, and Jialu Liu

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.696964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:b0c89a21ded79a014e37ba7403de5d56e0038faebfe129b1fc61aa0b2f621943

Observation ebebe5be-521a-4bce-bf94-a49137173aa8 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Reinforcement Learning for LLM Post-Training: A Survey Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:35:50.282923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:79ce42afc87eebc96df31d6e9bd12a0d15f3aac2fc6201e7a9c26df1567e2cb5

Observation 914132a8-46b6-498b-93ee-1e777a015ab2 · outbound

This paper cites Maas, Raymond E.

Reinforcement Learning for LLM Post-Training: A Survey Maas, Raymond E

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.689798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:2768829048a0e8911e2b13f7919c854046b4b63e36ac76bd250ac845cdbefc29

Observation 44b9fd79-d89e-4ca1-88e4-5c4a81c67f8f · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Reinforcement Learning for LLM Post-Training: A Survey Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:35:50.294514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:5dec7fa073c5faaa2db4c3f251123d621821f7385b00cca88db0e1e103a51ea1

Observation c127a630-88c7-4e30-9455-117dd01e7e21 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.

Reinforcement Learning for LLM Post-Training: A Survey Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.712970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:283cc7176908af386a0a4918b2116531d8d0f72e5141c6062d38d54d69a42132

Observation 6220a988-6323-437a-a3a2-8994ac19879a · outbound

This paper cites Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu.

Reinforcement Learning for LLM Post-Training: A Survey Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.652034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:636c65e64c5d54d5f4d23a5e92e09c9a334b08a2a237719cfdd4a5a7e0f542d7

Observation 93e86d94-962d-42d4-91ad-ad6bb5319b60 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Reinforcement Learning for LLM Post-Training: A Survey Xing, Hao Zhang, Joseph E

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.671441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:f28c9da2241ee7cb48b4f22533ac69120bbb2c280231b60118130b8e5b38f2af

Observation 45c694d0-ef68-4c86-b975-89963006acf4 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Reinforcement Learning for LLM Post-Training: A Survey Pythia: A suite for analyzing large language models across training and scaling

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.679386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:3e334f81c7852dd4efc3f21b4c0ee8e6865e970937d28714084824b10a465419

Observation af334596-b044-46de-a42a-9ed0357a83bd · outbound

This paper cites Solar 10.7b: Scaling large language models with simple yet effective depth up-scaling.

Reinforcement Learning for LLM Post-Training: A Survey Solar 10.7b: Scaling large language models with simple yet effective depth up-scaling

Reference 60

Resolution
malformed identifier
raw_fallback, observed 2026-05-23T22:35:51.619674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:f86ff47e5b978ee8815832b403f44fc1c7c856b76a660c735435b495dbe25b9e

Observation 7dffecc0-3606-4704-8e5d-56b776ea295d · outbound

This paper cites Orca: Progressive learning from complex explanation traces of gpt-4.

Reinforcement Learning for LLM Post-Training: A Survey Orca: Progressive learning from complex explanation traces of gpt-4

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.605131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:b407ef22bd2e76767fb890baddea8b4971246a2f07f5cc7473a6b3ad9513f8ab

Observation cbbc87f2-12e0-40f3-81a2-9b5a366cc81a · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback.

Reinforcement Learning for LLM Post-Training: A Survey Ultrafeedback: Boosting language models with high-quality feedback

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.590154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:896dc70a4f7c712d8faad65a11347bf2097e2311887a03f9ac464cb9230bf3d2

Observation 6cac6db5-fd88-47cc-8280-b276c08dc9aa · outbound

This paper cites Measuring massive multitask language understanding.

Reinforcement Learning for LLM Post-Training: A Survey Measuring massive multitask language understanding

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.593896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:58ce7c33414b4c446b8cd5166bcbbcdeeada7e29bb7d274105b1a938b3c6f9b5

Observation e9d142e5-78ac-4ab3-845a-ec23f0747172 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

Reinforcement Learning for LLM Post-Training: A Survey Winogrande: An adversarial winograd schema challenge at scale

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.609644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:6220ae04f23e81cf736b2794189c3fdff91281c46a83a6593b194ff93c87190f

Observation bf26ad90-ca31-454c-9f01-ca8cd4f79f12 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reinforcement Learning for LLM Post-Training: A Survey Training Verifiers to Solve Math Word Problems

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:35:50.288719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:0e0a7bac3cb8344a2846dec34dc21391c788e0ffe69e6dc1ee168a41e95d2885

Observation 737de695-d107-44c9-89d1-eac482791dbf · outbound

This paper cites Generalized preference optimization: A unified approach to offline alignment.

Reinforcement Learning for LLM Post-Training: A Survey Generalized preference optimization: A unified approach to offline alignment

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.570012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:d82b42853bda33ff04735811df55a5865f5203b59b55cdd0e2ae1d806cd04a91

Observation b3ea4c40-b1e4-47f9-9a82-ea6c2e86583e · outbound

This paper cites Language models are unsupervised multitask learners.

Reinforcement Learning for LLM Post-Training: A Survey Language models are unsupervised multitask learners

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.580449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:fade4c77d57c10ccda34cc066b6cad38c8b8cc9ef9d783601a4d92c2df02c33d

Observation 7bb92faf-99cb-49ac-a7d1-4fb093f39379 · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models.

Reinforcement Learning for LLM Post-Training: A Survey Llama 2: Open foundation and fine-tuned chat models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.587175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:2b31ac1be29d08eb86969436068dba5a0bb5ad8ed2e6a4191ee734d8f38d5c8c

Observation 829acdd5-5433-4a4f-bc7c-8e798de4b516 · outbound

This paper cites The cringe loss: Learning what language not to model.

Reinforcement Learning for LLM Post-Training: A Survey The cringe loss: Learning what language not to model

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.630359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:6ba48f95ffc065fbe26fb3c8d566db61eacc60439b91c1e18e2025c2febc0265

Observation d2973944-a7d3-47e5-adff-5e0316f35255 · outbound

This paper cites Hashimoto.

Reinforcement Learning for LLM Post-Training: A Survey Hashimoto

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.529270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:ecdfa6e3aab776dab3e43ec93a75270a796dd3e768f047749b3248a178478fe2

Observation 07ea3b54-dd38-4bb9-ab20-65e19538d55a · outbound

This paper cites Advances in prospect theory: Cumulative representation of uncertainty.

Reinforcement Learning for LLM Post-Training: A Survey Advances in prospect theory: Cumulative representation of uncertainty

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.526071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:1294d37d33b39e5d757698868e01c182ad9d9c16f41261e479b494b7cd7a56e1

Observation 66fabb02-424f-437d-9bf7-87525f861d18 · outbound

This paper cites Proximal policy optimization algorithms.

Reinforcement Learning for LLM Post-Training: A Survey Proximal policy optimization algorithms

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.532606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:8605baf21be7c1452e0cdbc13d5e97e5fd6b28a13bbf101624d63465a0b6d3fd

Observation 577b2d2d-1ae5-46ed-8d3f-1497d9cb6209 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.633995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:aaab1fee1438729f5b1c4c79fde8387743360413e81f60fffc3fea7268fe167f

Observation 07bff1ea-da89-472d-ad0b-613b07be387b · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Reinforcement Learning for LLM Post-Training: A Survey Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:35:50.299816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:6aac7a9a80c1a1880ba889c816dab4211331920b669086033cba8101d5aa2b12

Observation 1c6b788f-266b-496c-97aa-6397074e733c · outbound

This paper cites Phi-2: The surprising power of small language models.

Reinforcement Learning for LLM Post-Training: A Survey Phi-2: The surprising power of small language models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.563796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:f74e339a4667185ccb94454a3af3be1617cae5400eef866988331b97b4822134

Observation f7ef9e94-119e-42f1-a00e-468bdeafab0f · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.514946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:d76caa43ee548d5615ffe1941ba33eedba3590f94017f4e8892e3645b437f933

Observation daeb35c3-6af9-41c4-9bfc-1681de57d7b2 · outbound

This paper cites Instruction-following evaluation for large language models.

Reinforcement Learning for LLM Post-Training: A Survey Instruction-following evaluation for large language models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.518764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:3cd1bb4dd9a65d09a60bd4c701aab335a42d80656b07af34830ace36ce2d4f40

Observation a63075f8-1aa7-4c43-a0a0-a6a0848b2539 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Reinforcement Learning for LLM Post-Training: A Survey Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.545512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:3bc5e7789ee536a70f87b589865fd6d48f845d3eb3595194a6572ced23c37ded

Observation 17579027-2b4a-43c1-bcc2-7dcb0935f694 · outbound

This paper cites Gonzalez, and Ion Stoica.

Reinforcement Learning for LLM Post-Training: A Survey Gonzalez, and Ion Stoica

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.518033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:ee7e56081f65ec1a9abcedb35bbed8ab207556dfd6be3ba64ee3efe22a115158

Observation 3df9b7fd-5e9e-4321-a88c-f726d6c706ac · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

Reinforcement Learning for LLM Post-Training: A Survey Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.683292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:929eaea3331516ced1b21fa0e0af1450c76c9cba5632fc554037b87b78eb6a36

Observation 0d944cac-3c8d-46d5-a075-87e31d2ec18c · outbound

This paper cites Buy 4 REINFORCE samples, get a baseline for free!.

Reinforcement Learning for LLM Post-Training: A Survey Buy 4 REINFORCE samples, get a baseline for free!

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.552696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:b1a2d0cfe946f9055ebcdef9411b1b46edbc76d22a2904a81c2a069962ef047c

Observation 2db9056c-7483-4753-8b60-19302329c864 · outbound

This paper cites Learning to rank for information retrieval.

Reinforcement Learning for LLM Post-Training: A Survey Learning to rank for information retrieval

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.553126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:69f705e9f9923a04b95bbfc6998b005c17104b5e58334880d21be21d8e5b2a9e

Observation 5d0e78e9-0a98-4bc1-8f8f-ccdf8b8776a1 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.642809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:df6029779df44920865f1ad4cd67626860171d73624cd74ca876a47c64348b55

Observation b00bca99-5099-4c51-b2f5-c087e9a0f940 · outbound

This paper cites Openassistant conversations – democratizing large language model alignment.

Reinforcement Learning for LLM Post-Training: A Survey Openassistant conversations – democratizing large language model alignment

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.528599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:6ea997aff4c8d0c047f8a65630885c18be7c4e1c661fb15036ba94b9366d914e

Observation caa32368-772c-4fac-aa0e-8d37c2064eb6 · outbound

This paper cites Hashimoto.

Reinforcement Learning for LLM Post-Training: A Survey Hashimoto

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.559788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:3855a30eeea601424b23a895b4df3af02b14eb4c75c4bd319e9ce123707b40a4

Observation f41a71fb-eece-4c83-bb8b-db9a8deef6e5 · outbound

This paper cites Secrets of rlhf in large language models part ii: Reward modeling.

Reinforcement Learning for LLM Post-Training: A Survey Secrets of rlhf in large language models part ii: Reward modeling

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.571733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:5e6912602d77ba172ae4600ce38d3e9312c68480e7aa1315328ac839e0594765

Observation 732b60fd-bad0-4dd5-b241-e2d2243d738d · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

Reinforcement Learning for LLM Post-Training: A Survey Safe rlhf: Safe reinforcement learning from human feedback

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.657844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:0832799828b7d961666f533180b2191b1c2fee65257143f4ad2fcbde8f10d563

Observation a3466e86-ac7c-4a7d-a803-3cfa805308bf · outbound

This paper cites Lipton, and J.

Reinforcement Learning for LLM Post-Training: A Survey Lipton, and J

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.648279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:e60b916d663bc6f4c37831c4e8e36c48ab034a59bdb863cde26708e2cafacbd5

Observation 9ead4ea7-08ef-4c83-b1c6-b5622e5e4d4a · outbound

This paper cites A paradigm shift in machine translation: Boosting translation performance of large language models.

Reinforcement Learning for LLM Post-Training: A Survey A paradigm shift in machine translation: Boosting translation performance of large language models

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.701061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:e083721aa3ac4be149a4e35b71cf6e5fd569d5fae54ade1ab3c48a09bfa5c9a9

Observation 29ad9da9-968d-42cf-934d-f2077af82f7a · outbound

This paper cites On the limitations of the elo, real-world games are transitive, not additive.

Reinforcement Learning for LLM Post-Training: A Survey On the limitations of the elo, real-world games are transitive, not additive

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.705266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:8e68e9cf34af1772953551d5213a4261c793613157f519e2c9f3e26ce92eba32

Observation 35116a69-8d42-4971-b104-f0274d074a9c · outbound

This paper cites Self-play preference optimization for language model alignment.

Reinforcement Learning for LLM Post-Training: A Survey Self-play preference optimization for language model alignment

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.577157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:1e54e6a9b3a734adbda886dc0238a7bfa1d5cf4c5a462b54da8c23bcd51e3406

Observation e6375735-596d-49b2-b62b-63f99fa40c54 · outbound

This paper cites Schapire.

Reinforcement Learning for LLM Post-Training: A Survey Schapire

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.622936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:aae248a04942dba21c707b552475d14d0c46bdcfd078e4d402634bd0d3721a28

Observation e2eb508d-fac0-49e5-9c80-046ac23e41ea · outbound

This paper cites Llm-blender: Ensembling large language models with pairwise ranking and generative fusion.

Reinforcement Learning for LLM Post-Training: A Survey Llm-blender: Ensembling large language models with pairwise ranking and generative fusion

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.716609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:ac6cc7e56d77a1eb43fe9b0622172448856b47ba106e70def0b0fd8d8f80442e

Observation a58d8c4c-8b7f-4c89-8fbe-e47710ed9aa2 · outbound

This paper cites Orca 2: Teaching small language models how to reason.

Reinforcement Learning for LLM Post-Training: A Survey Orca 2: Teaching small language models how to reason

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.521463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:dabfe35ae32c1b76d62cf7f000a24edd0d12ea4b99560d3bfb9c6c9cdbc11044

Observation 10e9967c-47a8-4a89-a04e-dd069ca98ca6 · outbound

This paper cites On decoding strategies for neural text generators.

Reinforcement Learning for LLM Post-Training: A Survey On decoding strategies for neural text generators

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.501985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:0968dd7a8bc8d5341b3ef92c44760a89b3723e039d5e2aa6771b43af25b609e1

Observation 6892be44-c36d-4e92-b3c9-77c487ee1612 · outbound

This paper cites Insights into alignment: Evaluating dpo and its variants across multiple tasks.

Reinforcement Learning for LLM Post-Training: A Survey Insights into alignment: Evaluating dpo and its variants across multiple tasks

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.522881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:834de1074f8af1f7af8351907cc5e9781442704bcc63735f54aa6aaf3a0ede9a

Observation 422cf210-88ee-448d-9e2c-acbc1f761e69 · outbound

This paper cites Is dpo superior to ppo for llm alignment? a comprehensive study.

Reinforcement Learning for LLM Post-Training: A Survey Is dpo superior to ppo for llm alignment? a comprehensive study

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.563056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:f88335333e7c63173594fdf05e97febce66ef9eebb1db6dea8e513d5879ce579

Pith citing papers

Observation 3f4c7885-a7d7-4540-b853-7946b11876d2 · inbound

Language Models for Code Optimization: Survey, Challenges and Future Directions cites this paper.

Language Models for Code Optimization: Survey, Challenges and Future Directions Reinforcement Learning for LLM Post-Training: A Survey

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-10T22:34:34.886081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:34:34.886081Z digest=sha256:52a8ff0eefbb42c0737427c062acc9aa3418fe4c157e8c3c78e4ebe325c5317d

Observation e77c8d6b-4819-4097-b062-3685465a125b · inbound

Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information cites this paper.

Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information Reinforcement Learning for LLM Post-Training: A Survey

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:56.529090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:56.529090Z digest=sha256:8fe625fc070c79116f17621d368b98e0049d58afa28a21fa049464defe469547

Observation 97047068-fd81-44ab-b96e-4ee8ee8ddf25 · inbound

Normative Evaluation of Large Language Models with Everyday Moral Dilemmas cites this paper.

Normative Evaluation of Large Language Models with Everyday Moral Dilemmas Reinforcement Learning for LLM Post-Training: A Survey

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T00:49:57.120759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:49:57.120759Z digest=sha256:ea492e3a257d3f5514c258f510b8e8d9e22486d96700bf9e4c30e0fcc6157c16

Observation 36ea623a-0563-46d3-be3a-43262adaa47d · inbound

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment cites this paper.

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment Reinforcement Learning for LLM Post-Training: A Survey

Reference 223

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:12.669701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:12.669701Z digest=sha256:ae23c43683c0859f2ce6be1497b203f8f311e7ac0232177cd62a82195a19959c

Observation 473d809d-d83c-42f9-b01f-6af87fd5afa9 · inbound

Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception cites this paper.

Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception Reinforcement Learning for LLM Post-Training: A Survey

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T05:33:05.286073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:33:05.286073Z digest=sha256:8884c47585c9c4f0febe575f280525630f449609891f68162b4d9a41611cfa9b

Observation 2a2585f5-4521-4310-88dd-15dae3ba40d3 · inbound

HyPerAlign: Interpretable Personalized LLM Alignment via Hypothesis Generation cites this paper.

HyPerAlign: Interpretable Personalized LLM Alignment via Hypothesis Generation Reinforcement Learning for LLM Post-Training: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T05:18:07.033884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:18:07.033884Z digest=sha256:d967719358c6521f90ae23c486820b7631858d08bb34693b7ba091cbcb88b48a

Observation 89d9599f-d45c-407d-9ad9-74df626f7c6b · inbound

A Survey on Progress in LLM Alignment from the Perspective of Reward Design cites this paper.

A Survey on Progress in LLM Alignment from the Perspective of Reward Design Reinforcement Learning for LLM Post-Training: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:52:06.552166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:52:06.552166Z digest=sha256:2e5c1cdcb366a26fedae8fef3281d50ed2392c1506c3f253d66304ef45f8bc14

Observation cfef34eb-55a5-4923-b85b-7055b566bcf5 · inbound

Ethics and Persuasion in Reinforcement Learning from Human Feedback: A Procedural Rhetorical Approach cites this paper.

Ethics and Persuasion in Reinforcement Learning from Human Feedback: A Procedural Rhetorical Approach Reinforcement Learning for LLM Post-Training: A Survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:36.266375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:36.266375Z digest=sha256:309b68e2f9da31d82856f078159595ad42a5c21bffd3857cd252cbdfd0264af4

Observation b539ee6f-cebb-4b93-93e9-627b6d143fdb · inbound

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO cites this paper.

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Reinforcement Learning for LLM Post-Training: A Survey

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:09.406012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:09.406012Z digest=sha256:9b4c27422a4afcc401cf9997f5d1345d5a62d8955e07834bfbaaa2c231e64bc4

Observation 01410d90-fb2e-45c3-b15d-3c1fdd2cf1dd · inbound

Advancing LLM Safe Alignment with Safety Representation Ranking cites this paper.

Advancing LLM Safe Alignment with Safety Representation Ranking Reinforcement Learning for LLM Post-Training: A Survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.318170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.318170Z digest=sha256:29845907262fc4120654754fca0970effecac93d920449cc7e08522f7ba4c507

Observation ed33eb3c-588d-40f4-b5d4-4f68de34dea2 · inbound

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval cites this paper.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Reinforcement Learning for LLM Post-Training: A Survey

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.734701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.734701Z digest=sha256:c9bc008eb75d799d277153e561e48bfe8c162e294ca280714525c136f9149652

Observation fb5bcec9-0d8c-4b71-a4a5-f9d5840f3113 · inbound

Multi-Domain Explainability of Preferences cites this paper.

Multi-Domain Explainability of Preferences Reinforcement Learning for LLM Post-Training: A Survey

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:12.907001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:05:12.907001Z digest=sha256:3010421a84a62fac2ad2b9d29bfdf700d515fd7e1e2e4bde3f6de0beb6fcb5ec

Observation ac9e2e24-ee35-4fd7-b56a-a86ae8a39575 · inbound

Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities cites this paper.

Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities Reinforcement Learning for LLM Post-Training: A Survey

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:43:23.597483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:43:23.597483Z digest=sha256:10147d6ea8146410797df4acbd9195bd3221e7a12b2878d5af707b349e345d63

Observation b90d553c-2bc3-4b50-a060-c07269eedf3a · inbound

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction cites this paper.

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Reinforcement Learning for LLM Post-Training: A Survey

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T11:37:16.014516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T11:34:09.428653Z digest=sha256:8c1988fd6317c6b9170a00e05c913322821992844e91c19c1c8eb3af83ced2b9

Observation 22c6f668-2bb8-4441-afa9-2db16d208578 · inbound

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences cites this paper.

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences Reinforcement Learning for LLM Post-Training: A Survey

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:58.366027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:58.366027Z digest=sha256:1a92e8f03071ea213390184bd5134221d1eae63eeafc892c37976b5b2e8fd884

Observation d507ed32-0dc3-4897-99f5-69487eea9eca · inbound

Brevity is the soul of sustainability: Characterizing LLM response lengths cites this paper.

Brevity is the soul of sustainability: Characterizing LLM response lengths Reinforcement Learning for LLM Post-Training: A Survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:11.331229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:11.331229Z digest=sha256:dbe7a50e962a8f984629d3c45a396f36f32375b008f33dc9e5bf99393e5073b8

Observation 41321313-bb77-41b8-856f-84ad7fcbe144 · inbound

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model cites this paper.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Reinforcement Learning for LLM Post-Training: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.250044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.250044Z digest=sha256:1c33e15a2c5b0c139ce72dd5beb96101fb8d81a553dcfaa263e4d087acaea126

Observation 4bb4783a-b6b6-4dad-bb60-51766d8a175a · inbound

Exploring the Secondary Risks of Large Language Models cites this paper.

Exploring the Secondary Risks of Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T09:42:13.985529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T09:40:58.067398Z digest=sha256:0ac4b3768ce15f9b0786d9380dd48df539b1d1f27acf5c6845aabc13d5481c8a

Observation 48b5883c-1205-4056-9750-5aee35549958 · inbound

Simulating multiple human perspectives in socio-ecological systems using large language models cites this paper.

Simulating multiple human perspectives in socio-ecological systems using large language models Reinforcement Learning for LLM Post-Training: A Survey

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T18:22:43.299476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:22:43.299476Z digest=sha256:4db605d8b05f5929889c42d77e1e80e3c13085361fd6886008e88ea837910cf8

Observation ba916c9f-9c62-437c-9f14-46bee33a9acb · inbound

CTR-Driven Ad Text Generation via Online Feedback Preference Optimization cites this paper.

CTR-Driven Ad Text Generation via Online Feedback Preference Optimization Reinforcement Learning for LLM Post-Training: A Survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T13:48:46.717418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:48:46.717418Z digest=sha256:4beb76c7c73e38dff83d13e32d766faf1589a94aeb2aa354dcc2fe9d4cb57e42

Observation 6d5338f5-e98c-4264-b1fd-dbe29fbddc1b · inbound

An Explainable Emotion Alignment Framework for LLM-Empowered Agent in Metaverse Service Ecosystem cites this paper.

An Explainable Emotion Alignment Framework for LLM-Empowered Agent in Metaverse Service Ecosystem Reinforcement Learning for LLM Post-Training: A Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T11:53:17.754032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:53:17.754032Z digest=sha256:dae665e9e5c4beb231419ffb33ac90ce02eb20413f9d96b4b5617c6a1bb80ab2

Observation e3f38a03-98a7-4fcc-8e31-0856c86913a3 · inbound

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints cites this paper.

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints Reinforcement Learning for LLM Post-Training: A Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T21:33:14.784455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:33:14.784455Z digest=sha256:3e4d1ca26f4890152f69cf1b717a62e97631c8fe1c109fcb903f3946e1e52a3e

Observation 85a753f9-6bfc-45ea-99b9-89638003be12 · inbound

A Survey on Training-free Alignment of Large Language Models cites this paper.

A Survey on Training-free Alignment of Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T21:18:48.405513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:18:48.405513Z digest=sha256:7f06f8c98f657bc13455f96dda465a0b8214dcab2229896d24146354f2996fb1

Observation ebd692bf-faeb-40d8-9d2e-c6ff0bde5d3a · inbound

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention cites this paper.

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention Reinforcement Learning for LLM Post-Training: A Survey

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:56:51.816772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T20:54:30.449792Z digest=sha256:d7bd3dbe41b3b4e796e241df533f82588c0646f99d414187b11dba1804a089f0

Observation 5f5c0311-173e-4dcc-b796-221408854dff · inbound

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning cites this paper.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Reinforcement Learning for LLM Post-Training: A Survey

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T20:41:50.436344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:857eb34128e9abd58d2c76f03c6b2aaf3b7aca0539e7710802498f9cb7f5df77

Observation fd962b78-169e-4c63-8873-44cb2333d49d · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reinforcement Learning for LLM Post-Training: A Survey

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T19:21:48.376249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:db283a3da0f0d81422cfbd4019d16e46206d02a5858491c0d908d419648a9f2e

Observation cf9f655d-a984-48d5-a5e7-b7d55157347a · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 273

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:08.244079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:08.244079Z digest=sha256:3cbe707fba64114374f119d73e45a09903faeedd4926502d5fc30ab2408f079a

Observation e6b7f581-ab53-4d4e-a416-17ebd7aa179b · inbound

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial cites this paper.

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial Reinforcement Learning for LLM Post-Training: A Survey

Reference 265

Resolution
unresolved
no resolver link, observed 2026-08-05T04:50:32.523428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:50:32.523428Z digest=sha256:0d5a4544170a45a1a6b80669c2767f7591446ef4db2e994326bc7ba4d6a3ca14

Observation 2a15aba6-4d84-491b-bcb5-292e9565c697 · inbound

EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models cites this paper.

EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:03:16.345319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:03:16.345319Z digest=sha256:4337124a82230d2cf27c55ac137f577934deec6a139e2509090c95014aefc22d

Observation df2adc50-b14f-4bd0-b9ea-a23274a516e8 · inbound

Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey cites this paper.

Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey Reinforcement Learning for LLM Post-Training: A Survey

Reference 110

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:11:42.974325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T18:07:33.612858Z digest=sha256:5224ea0d739bfd908ebb08eed795ef28c9e49e04816c047f8a8e512b5d7e81ee

Observation f6300611-2eb1-4a5e-a01d-40ed2a7fe2ec · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reinforcement Learning for LLM Post-Training: A Survey

Reference 178

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:43.117616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:43.117616Z digest=sha256:fbf7f5b06d1c73d7a722d443058c7165cd97b89b75c368afbf69bf83ba94d8fe

Observation 2b16c4e7-1b7c-4756-b793-777bda1359ce · inbound

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training cites this paper.

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training Reinforcement Learning for LLM Post-Training: A Survey

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T20:30:35.472865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T20:30:17.581842Z digest=sha256:6d198e2860095f12203787f1639b9b482c55e663a1bed4bb55e9b8eac95e8e6f

Observation 1bfb6bdd-e510-45ea-b57f-d08bcbf58d1a · inbound

Voting with the Graph: Stable RLAIF via Topological Consistency Maximization cites this paper.

Voting with the Graph: Stable RLAIF via Topological Consistency Maximization Reinforcement Learning for LLM Post-Training: A Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T09:28:13.001684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:28:13.001684Z digest=sha256:ef4e38c99dad6463f66edf98a0e873324c308d5a281d12b559f6eb7b5a8e8b21

Observation 206d97fe-d9ff-4a62-b3a1-b31282853cbe · inbound

Maximizing the efficiency of human feedback in AI alignment: a comparative analysis cites this paper.

Maximizing the efficiency of human feedback in AI alignment: a comparative analysis Reinforcement Learning for LLM Post-Training: A Survey

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T22:02:16.789125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:02:16.789125Z digest=sha256:88aaf6c94251ab180c99910459006d2145815b2c71612f9ca02114419ceda3c4

Observation ee8de5cb-22a3-44e6-a3c6-92adee7be1f8 · inbound

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing cites this paper.

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing Reinforcement Learning for LLM Post-Training: A Survey

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T08:17:36.555589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T08:12:55.296932Z digest=sha256:0ae60bc51cac2686c02469b49e047b7f79e6b7186feb1da2fdb3ff8b90b5ff66

Observation 57d1d6d9-8bba-4db0-b197-cc43eed7a2ad · inbound

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection cites this paper.

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection Reinforcement Learning for LLM Post-Training: A Survey

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:40:41.730258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:131b0133ab64df35e58bab12ea193fd8728edcb9594af10ae434efcb46a8fa1e

Observation 92febfba-be07-4bce-956e-4459d2f6ea71 · inbound

VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models cites this paper.

VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:39:54.359255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T09:38:32.977517Z digest=sha256:2065620eac6670686b098ec0ad324fe602a7a3f98503bb9ab62f744a17c57d3a

Observation a87cf072-c53d-482d-9bf8-bd7207b29052 · inbound

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training cites this paper.

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training Reinforcement Learning for LLM Post-Training: A Survey

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T00:45:50.664592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:18:56.476698Z digest=sha256:04bf84a3bc0bb1337b2048d741cb4180b056612916ba7d0a3c8ad518a7baffed

Observation b479ae6a-3f9f-4648-b20e-7e70fce63411 · inbound

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization cites this paper.

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization Reinforcement Learning for LLM Post-Training: A Survey

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:11:01.113709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:11:49.149334Z digest=sha256:cd54753a6a3db5824e29cbb6079087ba852a7eaefbe9a16c7049c53f8badc384

Observation 07619aec-be4c-422f-9420-92393b1334c8 · inbound

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety cites this paper.

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety Reinforcement Learning for LLM Post-Training: A Survey

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:46:05.553273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T03:00:34.862711Z digest=sha256:69998338bf7eec387a16aa9ee3f5d375f07f64261af2f5169c8a26168a6f2b7c

Observation ad446046-50cc-4b8d-aaff-a5c5998ef8af · inbound

Pref-CTRL: Preference Driven LLM Alignment using Representation Editing cites this paper.

Pref-CTRL: Preference Driven LLM Alignment using Representation Editing Reinforcement Learning for LLM Post-Training: A Survey

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:19.342574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T06:28:14.378602Z digest=sha256:51c1528742c28821f4a078abde8697914196c0633862f2203bde38e3faee276d

Observation b5582d48-55c7-485b-9d8f-d0866a6e76b0 · inbound

Generating Place-Based Compromises Between Two Points of View cites this paper.

Generating Place-Based Compromises Between Two Points of View Reinforcement Learning for LLM Post-Training: A Survey

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:01:12.088276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T03:36:31.695964Z digest=sha256:adcebfbb186f497e3faf30c079d010bd9ae19bcdb10df69ab69f5a42dc29ef96

Observation 70210b30-9bb9-4993-a0ae-25a3c5d1a2e2 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:21:29.964715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:f48f6bb71d346d9c866e879c69f97f9f5e8bc9d2828bcfb67763720d7e196694

Observation 7d51e331-406f-4c13-9633-1e0b2752d58f · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:11:16.721393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:1adf68376470d44df2ea92a1b69a5c69a8ec753ca41be919e224f4ec16b1c54e

Observation 9e58f133-85e0-40f4-b5a1-a9d462a264cd · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:02:40.688757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:49649b2d92e04351903891282c3b73f70ba212733598f66f3538a8c04af4ceb4

Observation 8cb754df-3e9a-48c1-9039-97b4dcbdc9ff · inbound

Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance cites this paper.

Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance Reinforcement Learning for LLM Post-Training: A Survey

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T15:41:17.804578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-09T19:29:36.348680Z digest=sha256:0804248de23760a40d97706845e51e7c87d1b7dc72406d241d8ad434b9cdbe9d

Observation 2045f3d6-b30b-407c-adac-2243f11ef719 · inbound

Agents Should Replace Narrow Predictive AI as the Orchestrator in 6G AI-RAN cites this paper.

Agents Should Replace Narrow Predictive AI as the Orchestrator in 6G AI-RAN Reinforcement Learning for LLM Post-Training: A Survey

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T02:17:06.980568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T02:12:21.409833Z digest=sha256:e08f2f800341d9ca3dd0fb0519f2d184938aa7ba2ebc75282d38562d1a6a333d

Observation ebaa50f6-de7e-4137-8a96-7d14b955712b · inbound

UNIPO: Unified Interactive Visual Explanation for RL Fine-Tuning Policy Optimization cites this paper.

UNIPO: Unified Interactive Visual Explanation for RL Fine-Tuning Policy Optimization Reinforcement Learning for LLM Post-Training: A Survey

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:42:03.971815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T01:38:16.943754Z digest=sha256:6e4e6c2a91653cdcd69147c8725c8fd326dae58b66a69ca827f17691e2092855

Observation a5d940f6-bc0f-45d8-9b48-614becafa94a · inbound

Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning cites this paper.

Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning Reinforcement Learning for LLM Post-Training: A Survey

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T21:52:48.235735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T21:49:07.440832Z digest=sha256:91c7e0f9533b22b49dfd01f6e81d9ab5a0862de889264f76604cf331d9ffde8f

Observation 514183ef-2b4a-4335-a865-9861bc6f4aa6 · inbound

ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning cites this paper.

ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning Reinforcement Learning for LLM Post-Training: A Survey

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T22:53:49.325337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T22:51:56.666980Z digest=sha256:437e7a9da14d3a7ea3592b8da9c082b1f18a6ee28acfe9ffe76fcf0d757e3f54

Observation 1970e252-73e9-4025-abff-429e5b4c6713 · inbound

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment cites this paper.

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment Reinforcement Learning for LLM Post-Training: A Survey

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:06:56.031254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T02:28:54.243385Z digest=sha256:45bb7cd406f6ac26da522cf90b0c81fb970c7c9f76a67ce2c58b4b5b242c1e94

Observation f6851d24-0f0d-4dbb-9124-28228106628e · inbound

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation cites this paper.

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation Reinforcement Learning for LLM Post-Training: A Survey

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:34:26.678705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T08:30:37.016856Z digest=sha256:8829d77f17e4428ee560fcd3083663af3eae8c445fe73d48e1dabbb321a637f1

Observation a7acfd67-0f89-4f61-9688-2203b6f2ddcd · inbound

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support cites this paper.

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support Reinforcement Learning for LLM Post-Training: A Survey

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-07-01T12:35:43.844499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-01T01:57:54.065453Z digest=sha256:fe453ad3377da104e7664461c5a3e64e0f46b573ca8ce407160d0cb6b191155b

Observation 8da3d815-7164-449f-9afb-35f55a83bc06 · inbound

Meta-Learning Preferences for Multilingual LLM Alignment cites this paper.

Meta-Learning Preferences for Multilingual LLM Alignment Reinforcement Learning for LLM Post-Training: A Survey

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T05:36:16.410995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:36:16.410995Z digest=sha256:1823406a3a47892958ca9341959d42cb0ec80c513c40bef0a246a7a49de276ab

Observation 59a1edaa-924c-4831-a5be-ff2c11c0bace · inbound

Sound Probabilistic Safety Bounds for Large Language Models cites this paper.

Sound Probabilistic Safety Bounds for Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T10:22:41.813908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:22:41.813908Z digest=sha256:d5984ba380bb0ed565aa2076b34e2a40a3c38403ee90702f98d041477e69ef29

Observation d6cadd64-ae02-4da7-8972-71698e65181d · inbound

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment cites this paper.

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment Reinforcement Learning for LLM Post-Training: A Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-30T11:44:36.074744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:44:36.074744Z digest=sha256:1a0788ba850c4254a2f2af14b9540389e921d80cfb7b1c8cf9a0f4ee2dea029d

Observation fd7bca3c-28f0-40dd-ab7b-1af6eb095fa4 · inbound

Cleo: A Transparent and Controllable Chatbot for Conversational Commerce cites this paper.

Cleo: A Transparent and Controllable Chatbot for Conversational Commerce Reinforcement Learning for LLM Post-Training: A Survey

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:33.069090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:33.069090Z digest=sha256:5371826ade7531930337d43c8577a1484afa544bb8e45c63ef4f287f08ef7833

Observation f5153a63-dc2e-437c-9e86-d9e9f9d13f4f · inbound

Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning cites this paper.

Measure, Don't Optimize: Forecasting Recovery in LLM Unlearning Reinforcement Learning for LLM Post-Training: A Survey

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-15T14:17:57.024485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:17:57.024485Z digest=sha256:3ae401d0c3fc4fcd8f2b0765b9a5a2d496ad6c4ad266b81f6ea953649cbc9769