Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning for LLM Post-Training: A Survey

As of 5 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 35 inbound Pith citation observations for arXiv:2407.16216.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.16216 v4

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T22:35:35.287039Z

measured 131 of 131 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:39:08.244079Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

96 of 96 outbound references displayed

  • verified exact6
  • verified fuzzy81
  • unresolved8
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

7
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 41ee95c0-ed31-4565-8582-4bec1964d30a · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Reinforcement Learning for LLM Post-Training: A Survey Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.479044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:b8ccebb93671b0bea05515d9a050b9028c2d4fd36508edc753d432925c3582d9

Observation 5ab86491-1f11-4173-8582-3f0e9a83c074 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.482409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:c9b61bc2e5316abe02322aadbea28538935ec35cc2edacacf6042ce42554c6e1

Observation d6f3456d-b78f-48fb-91fe-f27f3bdc1e3b · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback.

Reinforcement Learning for LLM Post-Training: A Survey Training a helpful and harmless assistant with reinforcement learning from human feedback

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.486346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:5699fe56067351be73624671ed765277c844f8813d6c4d600c73b7b97cca6868

Observation 8bdeb67b-9621-411b-ac3d-f21b74ef9217 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.416492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:e97654bb590d50b0d430e7d7c75f19bc6a1c09b55addeac7b0151872030fe090

Observation a7510842-3eba-4869-9e23-ef019229ee74 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Reinforcement Learning for LLM Post-Training: A Survey The claude 3 model family: Opus, sonnet, haiku

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.475329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:ac4d547922adcbaf74f1975032ab922073755934e1dca05293121964f053e2ca

Observation 710a6eb1-3f02-4ac1-b880-ee18e4d8020b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Reinforcement Learning for LLM Post-Training: A Survey Gemini: A Family of Highly Capable Multimodal Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:35:50.304630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:7f27841a8ebb1044471be172e00d84ec76b03047ece9047a90015b7d27e0e284

Observation df9518d0-8156-496d-a784-1b0110a74a29 · outbound

This paper cites Rlhf workflow: From reward modeling to online rlhf.

Reinforcement Learning for LLM Post-Training: A Survey Rlhf workflow: From reward modeling to online rlhf

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.378072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:465918f01171509df4da082e1c228e282f89a730e094d206a2d646bc1b4d1f74

Observation 19cb9fd8-1e69-4193-8b94-52e65e2b02ce · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint.

Reinforcement Learning for LLM Post-Training: A Survey Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.397901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:48250262b24eb15d4d4f082f9879dce9a65ebe031a6060885240848a092bb2ac

Observation a5bedc3d-3223-4c19-9b81-8134f8cca61a · outbound

This paper cites Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan.

Reinforcement Learning for LLM Post-Training: A Survey Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.639341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:b3985626de36e9a12a29d67edcd0dbef8aaf589a009e2cdebd6651e3d7b19ee3

Observation 737657c7-ad21-433e-883a-e52a6140e94e · outbound

This paper cites Rlaif: Scaling reinforcement learning from human feedback with ai feedback.

Reinforcement Learning for LLM Post-Training: A Survey Rlaif: Scaling reinforcement learning from human feedback with ai feedback

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.573067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:0e6fa11bdb2f2832e6d96759da36ec138b50195e44da7aed551425bd0227f826

Observation 5bad53b0-48e8-417a-b203-af6f6f171569 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.583731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:10fc9c28431e13510c325c806b80668fa7a6ab3ede3f05a54f5860c755fd9128

Observation f9169425-faf9-488b-a6c6-bb694d11a966 · outbound

This paper cites Manning, and Chelsea Finn.

Reinforcement Learning for LLM Post-Training: A Survey Manning, and Chelsea Finn

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.545251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:e39171896891477321eca0c7a939396e0aa11a57757fc7e2648b368c8d636d13

Observation 90bfbf4d-0a77-4b8f-ac5e-338a80697b01 · outbound

This paper cites Smaug: Fixing failure modes of preference optimisation with dpo-positive.

Reinforcement Learning for LLM Post-Training: A Survey Smaug: Fixing failure modes of preference optimisation with dpo-positive

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.539904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:1666bfa440a6424ce1c63f274e80b0c3737cd9fd744e264d8fe5284981cd08a3

Observation 12833e43-3931-4d83-814e-a1eb31bcf212 · outbound

This paper cites β-dpo: Direct preference optimization with dynamic β.

Reinforcement Learning for LLM Post-Training: A Survey β-dpo: Direct preference optimization with dynamic β

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.549126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:48dbee8f479608d96cc6dcc38a3a1a1c87b47b42aef7c9815918d7776c103742

Observation 74acbde2-9219-46fb-bbbd-3f62361eb642 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Reinforcement Learning for LLM Post-Training: A Survey A general theoretical paradigm to understand learning from human preferences

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.565905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:9b3b73e5ca527a2044f4cf36f4bca97f849660d688ecec75b7a24f4946b83bb2

Observation ee29a6ea-2640-41c1-bd86-8919971c360d · outbound

This paper cites sdpo: Don’t use your data all at once.

Reinforcement Learning for LLM Post-Training: A Survey sdpo: Don’t use your data all at once

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.615981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:814009cd1b79eeb3a1361f53272f0611e563a0ef0bbd8a0f23b0d58bf063d04a

Observation f2cd99c0-c23d-4eae-bca1-fdd7cb9a4f4f · outbound

This paper cites From r to q∗: Your language model is secretly a q-function.

Reinforcement Learning for LLM Post-Training: A Survey From r to q∗: Your language model is secretly a q-function

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.693376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:94222939520122ad8f3027aecfaf56b86bfc451e97412425ba1dc21a07d025e0

Observation 37068e4e-f7d9-4b17-8516-1620c2cff80c · outbound

This paper cites Token-level direct preference optimization.

Reinforcement Learning for LLM Post-Training: A Survey Token-level direct preference optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.493725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:58ddde902e5cb5bf6fdabd28d52cb5d88f817b3584969bd840a73a6b0c435a41

Observation 2e864275-507d-4ae1-9998-22abf3fcbc60 · outbound

This paper cites Self-rewarding language models.

Reinforcement Learning for LLM Post-Training: A Survey Self-rewarding language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.470156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:faf08d34c75a015565967bea2c14f22b5e60187ae1454fe04e01e4c1b504c2a4

Observation 5ab8bc47-8da8-4d3e-99f2-450e8633bd52 · outbound

This paper cites Some things are more cringe than others: Iterative preference optimization with the pairwise cringe loss.

Reinforcement Learning for LLM Post-Training: A Survey Some things are more cringe than others: Iterative preference optimization with the pairwise cringe loss

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.498202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:76bca700ff6c708e9558af9fa64314e8183b42646601953a02de05e3699a82dc

Observation 14e27ea6-a84f-445e-8675-73cda460eed1 · outbound

This paper cites Kto: Model alignment as prospect theoretic optimization.

Reinforcement Learning for LLM Post-Training: A Survey Kto: Model alignment as prospect theoretic optimization

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.531924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:cca2f4b80487c50abc2c86116e59153bd2328f747f885299430a47ed03f887e5

Observation eb2c37f0-9d2b-40a5-bde0-424cc1cc69a9 · outbound

This paper cites Offline regularised reinforcement learning for large language models alignment.

Reinforcement Learning for LLM Post-Training: A Survey Offline regularised reinforcement learning for large language models alignment

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.514400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:8a2dfa1b2cbb3be470e2657daea62862d24ad334b493c7b08492f225909d6ac7

Observation 2dc7bb51-2f68-439e-bc2a-7eec04505a71 · outbound

This paper cites Orpo: Monolithic preference optimization without reference model.

Reinforcement Learning for LLM Post-Training: A Survey Orpo: Monolithic preference optimization without reference model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.507139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:18fde4cc8f3bab4a6f7f59bf7a9b1c038f25a8b98632f1b4399be7474503bc4c

Observation 3a15b19e-901b-4da0-895f-1ff4d2d85cad · outbound

This paper cites Paft: A parallel training paradigm for effective llm fine-tuning.

Reinforcement Learning for LLM Post-Training: A Survey Paft: A parallel training paradigm for effective llm fine-tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.462277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:fc89285aea6433418757aa463f69969b12bad584b585c3868ca9b5df2a34c7b8

Observation 2c05bf88-20f3-4fd2-854a-c08d4b8f7cc1 · outbound

This paper cites Disentangling length from quality in direct preference optimization.

Reinforcement Learning for LLM Post-Training: A Survey Disentangling length from quality in direct preference optimization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.489672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:3a8b41d608dad1b3244fc0e988dc1cad03006cc0737769dee18f9afe775c670b

Observation cba4f34d-36e6-4e27-8479-93be185c267e · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

Reinforcement Learning for LLM Post-Training: A Survey Simpo: Simple preference optimization with a reference-free reward

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.524829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:fe9ca5bd7ea58145aa40f6657cc65f3c4488de317d3a866e0c8b8577124d345a

Observation d7983a5b-7419-4b44-949b-68ce7060fac5 · outbound

This paper cites Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms.

Reinforcement Learning for LLM Post-Training: A Survey Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.441352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:4e82781681090625d19bd11505c13f42ef0bc198f305dd2e49ce52712e0af4be

Observation ff98a8ae-3b4e-4d9f-aaf1-44c676d78743 · outbound

This paper cites Liu, and Xuanhui Wang.

Reinforcement Learning for LLM Post-Training: A Survey Liu, and Xuanhui Wang

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.559480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:d370a32096baab6b60d992db546bbcb81f862ff69c0fcf03e86fddb5bdca1770

Observation d4d7fde4-792c-4b0d-b6a0-3cddc5a58f94 · outbound

This paper cites Rrhf: Rank responses to align language models with human feedback without tears.

Reinforcement Learning for LLM Post-Training: A Survey Rrhf: Rank responses to align language models with human feedback without tears

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.580163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:ff16cbe7d3c1af7b8e5e55444ab0b40c30acc03846d5ee897547c9beff546daf

Observation 11a31772-b1e4-4f5d-b2c1-99ac4f709805 · outbound

This paper cites Preference ranking optimization for human alignment.

Reinforcement Learning for LLM Post-Training: A Survey Preference ranking optimization for human alignment

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.503370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:dadd8c6377299161620bc862d9b3a05299fae9001167903850acd8bc8b2f6ef1

Observation e0b4a298-5c89-452c-b199-8df7ba70b1a6 · outbound

This paper cites Negating negatives: Alignment without human positive samples via distributional dispreference optimization.

Reinforcement Learning for LLM Post-Training: A Survey Negating negatives: Alignment without human positive samples via distributional dispreference optimization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.419438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:9e09b95037b85781a6c55a9520afbecbc9b6be5b1fe634213cfb07753d477607

Observation e143ba21-90f9-4991-871d-9f73a25da5c9 · outbound

This paper cites Negative preference optimization: From catastrophic collapse to effective unlearning.

Reinforcement Learning for LLM Post-Training: A Survey Negative preference optimization: From catastrophic collapse to effective unlearning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.511504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:9658c8a5b0807584f5f394018ad9ccac3dfb1d9b2ae4a640d901abba81a224c9

Observation aea6477c-9704-42a0-b7ed-bd91fc5afd24 · outbound

This paper cites Contrastive preference optimization: Pushing the boundaries of llm performance in machine translation.

Reinforcement Learning for LLM Post-Training: A Survey Contrastive preference optimization: Pushing the boundaries of llm performance in machine translation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.536251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:35a1373a7bb14c65357c85a5422ee6c4d2abb06f355d542ad1ee0b0256ebd6b0

Observation 81382c5c-effc-4216-bb54-8f9c4e55463e · outbound

This paper cites Mankowitz, Doina Precup, and Bilal Piot.

Reinforcement Learning for LLM Post-Training: A Survey Mankowitz, Doina Precup, and Bilal Piot

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.510618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:3503a10cbb5adad01ea16e2a076fbe4336c264c381870fd6cee863dc62e9a5d9

Observation a4f2e5ab-84a8-4f7b-8232-be7600e3cb30 · outbound

This paper cites A minimaximalist approach to reinforcement learning from human feedback.

Reinforcement Learning for LLM Post-Training: A Survey A minimaximalist approach to reinforcement learning from human feedback

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.415463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:6dbae8b7f2982853aae5b44730544b257cc93906472cd527b4ca6ce1533a6ad3

Observation d32da5e4-3c98-43c7-ba72-639be4d84958 · outbound

This paper cites Direct nash optimization: Teaching language models to self-improve with general preferences.

Reinforcement Learning for LLM Post-Training: A Survey Direct nash optimization: Teaching language models to self-improve with general preferences

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.401392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:471a6fecd3757e101b6805728049748f4ac1e06fc1ce58df5a7b5624041e34ee

Observation a8b485f9-bed1-48bb-b1a1-db9f3f5814d1 · outbound

This paper cites Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints.

Reinforcement Learning for LLM Post-Training: A Survey Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.408714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:ee47bf97c597239171e7aac90186bfd0de5e87a55423623936c946f108304391

Observation 7456c3d9-499d-4871-9292-2ecdbe5a8a2b · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.391681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:398e075dacb7ade5856c24f8f50ff7d50a804ad927f7b7c7a5b045a7e708bc7c

Observation a57386fa-1f0c-4055-976e-6be0bc487193 · outbound

This paper cites A markovian decision process.

Reinforcement Learning for LLM Post-Training: A Survey A markovian decision process

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.388440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:c2e4f58be7489f966f93ea191ee839a8668e01e99cc85f2292fd4eb8fc11c4c5

Observation 8ab25870-c9a6-46db-82e5-00fd6a328bc9 · outbound

This paper cites Hashimoto.

Reinforcement Learning for LLM Post-Training: A Survey Hashimoto

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.394727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:01ff15165fed5679321a905ae00b542fb5cd57d47eb318221c9946e0dc0cf54f

Observation a1778653-468c-4c57-ae85-77d0e4da3496 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Reinforcement Learning for LLM Post-Training: A Survey Bleu: a method for automatic evaluation of machine translation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.412117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:551cbdc20335b73509d4de332c062f9a3d7277b2a5a43a348d97a55adb77fce1

Observation 32f7872b-193f-4e3c-b8f5-2b2cb02ff5ce · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Reinforcement Learning for LLM Post-Training: A Survey Rouge: A package for automatic evaluation of summaries

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.482168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:3e20b0798e5c8340d4c6ae0895d5594c1e6c99b52e8555c11517745d3248eac8

Observation 7a2547ee-ce4c-427e-bdf1-793bd3cf4150 · outbound

This paper cites Weinberger, and Yoav Artzi.

Reinforcement Learning for LLM Post-Training: A Survey Weinberger, and Yoav Artzi

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.661237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:618b3f1a053c01a28bc78e9f09af801d854b023e5a544efc6acb7ee2f0370d79

Observation 83c152fe-d6fe-4aa6-8797-04e68279cce9 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.429877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:aa16329e1c59dfcae2e69302c978fd9b069d3627b10caba7766886e70cc99f86

Observation 696cf556-b69d-47cd-905b-9a70a9bd247f · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

Reinforcement Learning for LLM Post-Training: A Survey Truthfulqa: Measuring how models mimic human falsehoods

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.442972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:7ad8e3c0eb96d21ef7090a3f2dee1bcdc2bfd450ed97d3842ae1cd9a92b96b42

Observation 4d46e208-6801-4766-a3ec-de799503b78d · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Reinforcement Learning for LLM Post-Training: A Survey Chain-of-thought prompting elicits reasoning in large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.446993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:330eefd34b54f8c1e533b176a1d784195d58bc6772775bdf4cfa0793f30c8dea

Observation 611436f5-ae08-47ce-84a4-aa10272c94d3 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano.

Reinforcement Learning for LLM Post-Training: A Survey Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.665586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:eb8fe33ff23ebef860a54dd61739220047a4615fa1e4fc44de0144ad6f859c46

Observation b89e7f9e-2706-4d0b-b0f6-dac876fbcc1f · outbound

This paper cites Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H.

Reinforcement Learning for LLM Post-Training: A Survey Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.576240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:e41756e5216ab66a2c78bc16e30f7445f724a6e2348d8ad5121a0196fdd2e34e

Observation 48144c6b-97a2-42f1-a558-51791ce80eea · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Reinforcement Learning for LLM Post-Training: A Survey High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:35:50.309976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:651b7977f33a1813d09014ce2676702055695829e64fb69a1bb826442f605808

Observation 6e5b5ed8-83a9-4526-bca2-2fd5350c0a3b · outbound

This paper cites Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J.

Reinforcement Learning for LLM Post-Training: A Survey Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.444955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:88a86f806d046c8b02c1de8b1492675f5cec79871ca257678c2554ca00f9d2a4

Observation 985f1914-c302-4e1d-b75d-366da530af38 · outbound

This paper cites Liu, and Jialu Liu.

Reinforcement Learning for LLM Post-Training: A Survey Liu, and Jialu Liu

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.696964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:0b3fd819b27f504bc5a912cd990afd7cf02a54ea9a04973566861aa28d0be722

Observation ebebe5be-521a-4bce-bf94-a49137173aa8 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Reinforcement Learning for LLM Post-Training: A Survey Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:35:50.282923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:d61ce0e5a95e8c21a67138328c8cadb69ec339af9d0b526e2750bbed7398d53b

Observation 914132a8-46b6-498b-93ee-1e777a015ab2 · outbound

This paper cites Maas, Raymond E.

Reinforcement Learning for LLM Post-Training: A Survey Maas, Raymond E

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.689798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:a914e8b2c87f79b8028e2fcbdf75ab67277fe0d16cc90f95531493a8ae176e0d

Observation 44b9fd79-d89e-4ca1-88e4-5c4a81c67f8f · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Reinforcement Learning for LLM Post-Training: A Survey Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:35:50.294514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:504122048da6331e63d4af3f5f372eb47792cb914bdb29c51835778a711fccc8

Observation c127a630-88c7-4e30-9455-117dd01e7e21 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.

Reinforcement Learning for LLM Post-Training: A Survey Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.712970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:b3fe56d2e55edf093b3f61253e228ef059f5d61ae3d0306d9924caf1d6b6fcf3

Observation 6220a988-6323-437a-a3a2-8994ac19879a · outbound

This paper cites Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu.

Reinforcement Learning for LLM Post-Training: A Survey Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.652034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:2d3dd6b1403b5a0a393b2d71e9f48584eac82eb5efc1af707f30e155edb1cb18

Observation 93e86d94-962d-42d4-91ad-ad6bb5319b60 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Reinforcement Learning for LLM Post-Training: A Survey Xing, Hao Zhang, Joseph E

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.671441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:e0b5f1b7365d8ab12a9aa1908b13714a70a2c16c6addb67573008c4dfec897e0

Observation 45c694d0-ef68-4c86-b975-89963006acf4 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Reinforcement Learning for LLM Post-Training: A Survey Pythia: A suite for analyzing large language models across training and scaling

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.679386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:8dee046999cafe14c5f3c5503b60309f6cd81688d4d2cacdbadc2fd33a9a9784

Observation af334596-b044-46de-a42a-9ed0357a83bd · outbound

This paper cites Solar 10.7b: Scaling large language models with simple yet effective depth up-scaling.

Reinforcement Learning for LLM Post-Training: A Survey Solar 10.7b: Scaling large language models with simple yet effective depth up-scaling

Reference 60

Resolution
malformed identifier
raw_fallback, observed 2026-05-23T22:35:51.619674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:bb95941b0d187f3a013fc7b9dcaea2b8dd6d8e4150afb7d9e87c88538e0cbd05

Observation 7dffecc0-3606-4704-8e5d-56b776ea295d · outbound

This paper cites Orca: Progressive learning from complex explanation traces of gpt-4.

Reinforcement Learning for LLM Post-Training: A Survey Orca: Progressive learning from complex explanation traces of gpt-4

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.605131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:a611e686e0280a6955227123d4768ec94c35a1ca7e39d9c67ac7497b39633fb3

Observation cbbc87f2-12e0-40f3-81a2-9b5a366cc81a · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback.

Reinforcement Learning for LLM Post-Training: A Survey Ultrafeedback: Boosting language models with high-quality feedback

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.590154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:5632da0ce5f6404a034077715d8121e46b662fb5bb06a3d62b0d3e3ba6d9a8f8

Observation 6cac6db5-fd88-47cc-8280-b276c08dc9aa · outbound

This paper cites Measuring massive multitask language understanding.

Reinforcement Learning for LLM Post-Training: A Survey Measuring massive multitask language understanding

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.593896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:d2d69661ab5e3fc3f9880570d6a2a551d86abc2c186cda5e77cbe49cc6fe658f

Observation e9d142e5-78ac-4ab3-845a-ec23f0747172 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

Reinforcement Learning for LLM Post-Training: A Survey Winogrande: An adversarial winograd schema challenge at scale

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.609644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:8c902dffe1ddf58d005760d9e784d069354312f555c23cf6d9844799119bbee9

Observation bf26ad90-ca31-454c-9f01-ca8cd4f79f12 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reinforcement Learning for LLM Post-Training: A Survey Training Verifiers to Solve Math Word Problems

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:35:50.288719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:a881da14d7bd002b6e93a453c1b49d387ead20372f6e29b945ec2e125889241e

Observation 737de695-d107-44c9-89d1-eac482791dbf · outbound

This paper cites Generalized preference optimization: A unified approach to offline alignment.

Reinforcement Learning for LLM Post-Training: A Survey Generalized preference optimization: A unified approach to offline alignment

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.570012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:960a04ea50205fef1016f6b3a2fbe48d44f904286772df020da160a27f84baa5

Observation b3ea4c40-b1e4-47f9-9a82-ea6c2e86583e · outbound

This paper cites Language models are unsupervised multitask learners.

Reinforcement Learning for LLM Post-Training: A Survey Language models are unsupervised multitask learners

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.580449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:fe586571e1662404aab6cf68fd9a0a2b31755c013a6fb26e56a6d438cfc4bc7d

Observation 7bb92faf-99cb-49ac-a7d1-4fb093f39379 · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models.

Reinforcement Learning for LLM Post-Training: A Survey Llama 2: Open foundation and fine-tuned chat models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.587175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:d59706237cf90ff780cabd0bb98e13f636a30357f8f09d72ba7d2d2562cb1e82

Observation 829acdd5-5433-4a4f-bc7c-8e798de4b516 · outbound

This paper cites The cringe loss: Learning what language not to model.

Reinforcement Learning for LLM Post-Training: A Survey The cringe loss: Learning what language not to model

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.630359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:27ad0bc6d1538f110997ce5ae5dab6ecc93b677bd6b7828a7ef1f49dc22c9529

Observation d2973944-a7d3-47e5-adff-5e0316f35255 · outbound

This paper cites Hashimoto.

Reinforcement Learning for LLM Post-Training: A Survey Hashimoto

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.529270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:bb15f17bb3e7fd13ce9acf09a5919a56349c4457a5a4cc04c7e7da2f43b78caa

Observation 07ea3b54-dd38-4bb9-ab20-65e19538d55a · outbound

This paper cites Advances in prospect theory: Cumulative representation of uncertainty.

Reinforcement Learning for LLM Post-Training: A Survey Advances in prospect theory: Cumulative representation of uncertainty

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.526071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:3ccb8915826626f7a75f7fdbd4e301e7258386177bf92d18a8ef6e3b338a9079

Observation 66fabb02-424f-437d-9bf7-87525f861d18 · outbound

This paper cites Proximal policy optimization algorithms.

Reinforcement Learning for LLM Post-Training: A Survey Proximal policy optimization algorithms

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.532606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:b6489fd28043026aedd0d412b40c7a1c31e3ac74e1fc2c7ea7e09525112f7c5b

Observation 577b2d2d-1ae5-46ed-8d3f-1497d9cb6209 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.633995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:d7ad552f1f12c7fd7f68d3735452afc360219679e3a4b0115f626bc2a388993a

Observation 07bff1ea-da89-472d-ad0b-613b07be387b · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Reinforcement Learning for LLM Post-Training: A Survey Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:35:50.299816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:4d7c12ad356fbd75a3b58bb16a5fd18ee9057134faec795958a1a28fd9818fa6

Observation 1c6b788f-266b-496c-97aa-6397074e733c · outbound

This paper cites Phi-2: The surprising power of small language models.

Reinforcement Learning for LLM Post-Training: A Survey Phi-2: The surprising power of small language models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.563796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:f37144d3f9676ea3f17bb1dddd23eab3bb48a6053f3b310a3c93e66c4a426e04

Observation f7ef9e94-119e-42f1-a00e-468bdeafab0f · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.514946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:7c5ac8e7a92a960501912af0574218bfbdb1e17068ac10632eff4c08cc2728a4

Observation daeb35c3-6af9-41c4-9bfc-1681de57d7b2 · outbound

This paper cites Instruction-following evaluation for large language models.

Reinforcement Learning for LLM Post-Training: A Survey Instruction-following evaluation for large language models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.518764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:be740918740490f535d2e09d4bd710dbd1b42ff9e89bcbdb4b2aa91c63fdcbfb

Observation a63075f8-1aa7-4c43-a0a0-a6a0848b2539 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Reinforcement Learning for LLM Post-Training: A Survey Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.545512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:48495eaa56bbcf8b02bcbd99e689ac7e6d88f2e40cee34b275e89922aad6d586

Observation 17579027-2b4a-43c1-bcc2-7dcb0935f694 · outbound

This paper cites Gonzalez, and Ion Stoica.

Reinforcement Learning for LLM Post-Training: A Survey Gonzalez, and Ion Stoica

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.518033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:60f3943d9e58a72eb4ca57bc65627971da60fd500916ae4b676dee6b8e86da60

Observation 3df9b7fd-5e9e-4321-a88c-f726d6c706ac · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

Reinforcement Learning for LLM Post-Training: A Survey Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.683292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:0b3d29cf6da264dc37f3e097272bb44758adf7cf2763d6e54f335bbb12d601f0

Observation 0d944cac-3c8d-46d5-a075-87e31d2ec18c · outbound

This paper cites Buy 4 REINFORCE samples, get a baseline for free!.

Reinforcement Learning for LLM Post-Training: A Survey Buy 4 REINFORCE samples, get a baseline for free!

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.552696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:5763f41e6e4817c02c54715187ca78ab74aa0be323129400cece6a581b14a6a3

Observation 2db9056c-7483-4753-8b60-19302329c864 · outbound

This paper cites Learning to rank for information retrieval.

Reinforcement Learning for LLM Post-Training: A Survey Learning to rank for information retrieval

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.553126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:df72e0c16abb858a66c5848268589f77f4f3373a9a646f4a2df3d6b6612e1d11

Observation 5d0e78e9-0a98-4bc1-8f8f-ccdf8b8776a1 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning for LLM Post-Training: A Survey Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-05-23T22:35:51.642809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:3dab42d96f3785eeacf532c5e84e58596d1cbbcceaa51cff626772d4287f0b7e

Observation b00bca99-5099-4c51-b2f5-c087e9a0f940 · outbound

This paper cites Openassistant conversations – democratizing large language model alignment.

Reinforcement Learning for LLM Post-Training: A Survey Openassistant conversations – democratizing large language model alignment

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.528599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:042bd17817235935e68df093c092390439c09dd34abe5cabd85f5eb26b059cbd

Observation caa32368-772c-4fac-aa0e-8d37c2064eb6 · outbound

This paper cites Hashimoto.

Reinforcement Learning for LLM Post-Training: A Survey Hashimoto

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.559788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:7f759b77157bef9e09c7f931d30f2c863bda5378e5b9a4a93970d6313011a0e2

Observation f41a71fb-eece-4c83-bb8b-db9a8deef6e5 · outbound

This paper cites Secrets of rlhf in large language models part ii: Reward modeling.

Reinforcement Learning for LLM Post-Training: A Survey Secrets of rlhf in large language models part ii: Reward modeling

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.571733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:3e9ef42258006ebd8de02863d6cca1e5acb1f42f935ce780fa071f3cf1f1daca

Observation 732b60fd-bad0-4dd5-b241-e2d2243d738d · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

Reinforcement Learning for LLM Post-Training: A Survey Safe rlhf: Safe reinforcement learning from human feedback

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.657844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:3fec352fbbd6bad1d7046738cc7bf56c48c181c8df19eb8ad8d83c383b107d0f

Observation a3466e86-ac7c-4a7d-a803-3cfa805308bf · outbound

This paper cites Lipton, and J.

Reinforcement Learning for LLM Post-Training: A Survey Lipton, and J

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.648279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:d1b73f4377134115a2730fb178ced0c23865bcfd547619b7522607b076e07e39

Observation 9ead4ea7-08ef-4c83-b1c6-b5622e5e4d4a · outbound

This paper cites A paradigm shift in machine translation: Boosting translation performance of large language models.

Reinforcement Learning for LLM Post-Training: A Survey A paradigm shift in machine translation: Boosting translation performance of large language models

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.701061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:54752e7bdebf3d9450d795e0ca03398824ee8d9335f487722cb26e88be20bf1a

Observation 29ad9da9-968d-42cf-934d-f2077af82f7a · outbound

This paper cites On the limitations of the elo, real-world games are transitive, not additive.

Reinforcement Learning for LLM Post-Training: A Survey On the limitations of the elo, real-world games are transitive, not additive

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.705266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:ae4a2d084b9a616ae700f1e63c24e9ff275e9c1a4b316b41dd5e16f2b313e08a

Observation 35116a69-8d42-4971-b104-f0274d074a9c · outbound

This paper cites Self-play preference optimization for language model alignment.

Reinforcement Learning for LLM Post-Training: A Survey Self-play preference optimization for language model alignment

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.577157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:a93aab949fbcb3323def3e490df91f7294774fe0e0e53ddea26e7183581a339a

Observation e6375735-596d-49b2-b62b-63f99fa40c54 · outbound

This paper cites Schapire.

Reinforcement Learning for LLM Post-Training: A Survey Schapire

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.622936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:c6a3aeb4a6947742501742512255bde9b603ba379fa2e4e13e0f6e182c35f73e

Observation e2eb508d-fac0-49e5-9c80-046ac23e41ea · outbound

This paper cites Llm-blender: Ensembling large language models with pairwise ranking and generative fusion.

Reinforcement Learning for LLM Post-Training: A Survey Llm-blender: Ensembling large language models with pairwise ranking and generative fusion

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.716609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:e00072d0cd3a004458a2b1acc85b691140d892f3f8d9db7110e1965d053bba0a

Observation a58d8c4c-8b7f-4c89-8fbe-e47710ed9aa2 · outbound

This paper cites Orca 2: Teaching small language models how to reason.

Reinforcement Learning for LLM Post-Training: A Survey Orca 2: Teaching small language models how to reason

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.521463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:f1615b6c664dc7658f5aec8cc64dd7bc3d9dd7736fadab08a1b4d871e5cff8aa

Observation 10e9967c-47a8-4a89-a04e-dd069ca98ca6 · outbound

This paper cites On decoding strategies for neural text generators.

Reinforcement Learning for LLM Post-Training: A Survey On decoding strategies for neural text generators

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.501985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:2815f1f4d0076f585dcd5f6a35487cb41663957b1623b64e5d353b5ea51d6c12

Observation 6892be44-c36d-4e92-b3c9-77c487ee1612 · outbound

This paper cites Insights into alignment: Evaluating dpo and its variants across multiple tasks.

Reinforcement Learning for LLM Post-Training: A Survey Insights into alignment: Evaluating dpo and its variants across multiple tasks

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.522881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:b9a8bc69ef4ab024ac5632e946c88af85e30783724c8df09dc88f60d99ce2f09

Observation 422cf210-88ee-448d-9e2c-acbc1f761e69 · outbound

This paper cites Is dpo superior to ppo for llm alignment? a comprehensive study.

Reinforcement Learning for LLM Post-Training: A Survey Is dpo superior to ppo for llm alignment? a comprehensive study

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T22:35:51.563056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:e71cd05a9b2f3563ce16a92b234474d91e970eea38c285e2e441c690519e7b30

Pith citing papers

Observation b90d553c-2bc3-4b50-a060-c07269eedf3a · inbound

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction cites this paper.

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Reinforcement Learning for LLM Post-Training: A Survey

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T11:37:16.014516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T11:34:09.428653Z digest=sha256:ff45f77690c9d217f34d05584d1a5c8cab6def3e41739b831ddbeb144e988f6a

Observation 4bb4783a-b6b6-4dad-bb60-51766d8a175a · inbound

Exploring the Secondary Risks of Large Language Models cites this paper.

Exploring the Secondary Risks of Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T09:42:13.985529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T09:40:58.067398Z digest=sha256:56086211ec8ed68fd9931f64b3fe0450f967cfd6b56f3c555a8234c192014c11

Observation ebd692bf-faeb-40d8-9d2e-c6ff0bde5d3a · inbound

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention cites this paper.

EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention Reinforcement Learning for LLM Post-Training: A Survey

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-18T20:56:51.816772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T20:54:30.449792Z digest=sha256:d22902ff3360925860a716b2a150f5a861975a139c44774a270f3f81819d6402

Observation 5f5c0311-173e-4dcc-b796-221408854dff · inbound

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning cites this paper.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Reinforcement Learning for LLM Post-Training: A Survey

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T20:41:50.436344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:17c7de85b86942d6fac8f9b18270170d0b8a3b53ea926f49a18484d7f308148e

Observation fd962b78-169e-4c63-8873-44cb2333d49d · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Reinforcement Learning for LLM Post-Training: A Survey

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T19:21:48.376249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:555bb71e2d54402d74eaaae0bd1870f091d109438bac843f689df588a5aba968

Observation cf9f655d-a984-48d5-a5e7-b7d55157347a · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 273

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:08.244079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:08.244079Z digest=sha256:432126b28831c8e04b959006e1d67007befcc1195d916d811b04ce73bf8b962b

Observation e6b7f581-ab53-4d4e-a416-17ebd7aa179b · inbound

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial cites this paper.

Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial Reinforcement Learning for LLM Post-Training: A Survey

Reference 265

Resolution
unresolved
no resolver link, observed 2026-08-05T04:50:32.523428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:50:32.523428Z digest=sha256:775b7dfae91a2e40d51a3d2186a5c6e433b9c9903dabe9d5f1498ac755d0aa8d

Observation 2a15aba6-4d84-491b-bcb5-292e9565c697 · inbound

EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models cites this paper.

EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:03:16.345319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:03:16.345319Z digest=sha256:d42e3b52c2a005f131dd5a998f54a03a26b1d10339d771c07d78aedeef99e27e

Observation df2adc50-b14f-4bd0-b9ea-a23274a516e8 · inbound

Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey cites this paper.

Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey Reinforcement Learning for LLM Post-Training: A Survey

Reference 110

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:11:42.974325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T18:07:33.612858Z digest=sha256:97ca77a11361c0424072b3cdda7991dbba6ff6ecd44148d0668ca1e754eb9e51

Observation f6300611-2eb1-4a5e-a01d-40ed2a7fe2ec · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reinforcement Learning for LLM Post-Training: A Survey

Reference 178

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:43.117616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:43.117616Z digest=sha256:dd376f46dad2b6d0ce1a3f59ca78aca95128fe04b06533bc4b7a8c1c221da0c2

Observation 2b16c4e7-1b7c-4756-b793-777bda1359ce · inbound

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training cites this paper.

SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training Reinforcement Learning for LLM Post-Training: A Survey

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T20:30:35.472865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T20:30:17.581842Z digest=sha256:4d7cb3a9d90ea3818b8f66ff61a98e7750c37bca4c897b627784afd2b67bea82

Observation 1bfb6bdd-e510-45ea-b57f-d08bcbf58d1a · inbound

Voting with the Graph: Stable RLAIF via Topological Consistency Maximization cites this paper.

Voting with the Graph: Stable RLAIF via Topological Consistency Maximization Reinforcement Learning for LLM Post-Training: A Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T09:28:13.001684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:28:13.001684Z digest=sha256:df9b1839bedb516f8eb04220e1fb0589aabc69f2e441d572c67dbe1c8fd2bc36

Observation 206d97fe-d9ff-4a62-b3a1-b31282853cbe · inbound

Maximizing the efficiency of human feedback in AI alignment: a comparative analysis cites this paper.

Maximizing the efficiency of human feedback in AI alignment: a comparative analysis Reinforcement Learning for LLM Post-Training: A Survey

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T22:02:16.789125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:02:16.789125Z digest=sha256:398035bd67af062e7e5f0b34f8e66446037e4a154eb363b8951da08ec93b8a59

Observation ee8de5cb-22a3-44e6-a3c6-92adee7be1f8 · inbound

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing cites this paper.

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing Reinforcement Learning for LLM Post-Training: A Survey

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T08:17:36.555589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T08:12:55.296932Z digest=sha256:248d223b1f86a573bb9ad0b4dd70558bfdeccfd6948b01b9ad76008b310ae3dc

Observation 57d1d6d9-8bba-4db0-b197-cc43eed7a2ad · inbound

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection cites this paper.

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection Reinforcement Learning for LLM Post-Training: A Survey

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:40:41.730258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:5fa8300c679da1c9f41ec2c46bf995a92c04522e496717d5dd01fa6402877bfa

Observation 92febfba-be07-4bce-956e-4459d2f6ea71 · inbound

VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models cites this paper.

VC-Soup: Value-Consistency Guided Multi-Value Alignment for Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:39:54.359255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T09:38:32.977517Z digest=sha256:5a6d65757c2f2ab215f835eca177dede982a52d0c0450be51cf41ac713813474

Observation a87cf072-c53d-482d-9bf8-bd7207b29052 · inbound

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training cites this paper.

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training Reinforcement Learning for LLM Post-Training: A Survey

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T00:45:50.664592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:18:56.476698Z digest=sha256:979f278c54b6d4ca29cafb9d6bdb646d88c47394775668460f96e9d921b21656

Observation b479ae6a-3f9f-4648-b20e-7e70fce63411 · inbound

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization cites this paper.

Mobile GUI Agent Privacy Personalization with Trajectory Induced Preference Optimization Reinforcement Learning for LLM Post-Training: A Survey

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:11:01.113709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:11:49.149334Z digest=sha256:c2efb14cf51dc5a1b6f43f6fb8765b40554628989f12af5ff03e6cb08ed4e9cc

Observation 07619aec-be4c-422f-9420-92393b1334c8 · inbound

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety cites this paper.

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety Reinforcement Learning for LLM Post-Training: A Survey

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:46:05.553273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T03:00:34.862711Z digest=sha256:4f2c1f65bf818f646a77a5452d2ee4d02fd067bc6cd7567a3546dbe062b448e9

Observation ad446046-50cc-4b8d-aaff-a5c5998ef8af · inbound

Pref-CTRL: Preference Driven LLM Alignment using Representation Editing cites this paper.

Pref-CTRL: Preference Driven LLM Alignment using Representation Editing Reinforcement Learning for LLM Post-Training: A Survey

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:11:19.342574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T06:28:14.378602Z digest=sha256:465fc415d0f9f080610f989863d76f37529a4b98e9037a91454bfa983d0d26ae

Observation b5582d48-55c7-485b-9d8f-d0866a6e76b0 · inbound

Generating Place-Based Compromises Between Two Points of View cites this paper.

Generating Place-Based Compromises Between Two Points of View Reinforcement Learning for LLM Post-Training: A Survey

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:01:12.088276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T03:36:31.695964Z digest=sha256:ed2627a1ce56ecd431d07242bece8dd505ebfbcdb9a31ce8694864426a37cc49

Observation 70210b30-9bb9-4993-a0ae-25a3c5d1a2e2 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:21:29.964715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:674551d6ee40fd9084816387712e600ccd3d6598555b78b6c52099e194ed18e2

Observation 7d51e331-406f-4c13-9633-1e0b2752d58f · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:11:16.721393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:4319c247061614cdc4a992dcaa76307d20df3a514eb711d571d3fcae3d115558

Observation 9e58f133-85e0-40f4-b5a1-a9d462a264cd · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:02:40.688757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:b5a8a96081b827263b5e86077b3f94e3f91859970bab9b954268d48004c032ed

Observation 8cb754df-3e9a-48c1-9039-97b4dcbdc9ff · inbound

Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance cites this paper.

Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance Reinforcement Learning for LLM Post-Training: A Survey

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T15:41:17.804578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T19:29:36.348680Z digest=sha256:b1678fcf79cc75d26c7a5bfd9a33b53bfb42de027e5b164529a8f9b99bb03e9b

Observation 2045f3d6-b30b-407c-adac-2243f11ef719 · inbound

Agents Should Replace Narrow Predictive AI as the Orchestrator in 6G AI-RAN cites this paper.

Agents Should Replace Narrow Predictive AI as the Orchestrator in 6G AI-RAN Reinforcement Learning for LLM Post-Training: A Survey

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T02:17:06.980568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:12:21.409833Z digest=sha256:55fdc0d506da6e9cbf61ec210c8eeb47e71ab64d85bc702b5388c47f025717c2

Observation ebaa50f6-de7e-4137-8a96-7d14b955712b · inbound

UNIPO: Unified Interactive Visual Explanation for RL Fine-Tuning Policy Optimization cites this paper.

UNIPO: Unified Interactive Visual Explanation for RL Fine-Tuning Policy Optimization Reinforcement Learning for LLM Post-Training: A Survey

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:42:03.971815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:38:16.943754Z digest=sha256:0210e98d7b5f6b9ba90d800ea944aabb4eb3f6b97a098afc429da9faadca1636

Observation a5d940f6-bc0f-45d8-9b48-614becafa94a · inbound

Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning cites this paper.

Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning Reinforcement Learning for LLM Post-Training: A Survey

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T21:52:48.235735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T21:49:07.440832Z digest=sha256:26a2d293ffb1306baa640b5e82ef54247ffa7ce72d2bfeb47648539b6bd5da70

Observation 514183ef-2b4a-4335-a865-9861bc6f4aa6 · inbound

ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning cites this paper.

ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning Reinforcement Learning for LLM Post-Training: A Survey

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T22:53:49.325337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T22:51:56.666980Z digest=sha256:db74d89449ffd4dd49b2790d472831b68b7d73015b337d78f2d235a761fc8946

Observation 1970e252-73e9-4025-abff-429e5b4c6713 · inbound

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment cites this paper.

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment Reinforcement Learning for LLM Post-Training: A Survey

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:06:56.031254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T02:28:54.243385Z digest=sha256:0e1a9e5b1e5e3f1ca7c8811218a96423cfc02350855c466f4b64219a2499fa30

Observation f6851d24-0f0d-4dbb-9124-28228106628e · inbound

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation cites this paper.

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation Reinforcement Learning for LLM Post-Training: A Survey

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:34:26.678705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T08:30:37.016856Z digest=sha256:7de33ab026cb5fa403cce75c728335d238789b3f85066f202862a400fa47b588

Observation a7acfd67-0f89-4f61-9688-2203b6f2ddcd · inbound

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support cites this paper.

Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support Reinforcement Learning for LLM Post-Training: A Survey

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-07-01T12:35:43.844499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-01T01:57:54.065453Z digest=sha256:2f28bb5cda3f4e05c919cff96e463edfe18ac62381a270506634f82935eabdaa

Observation 8da3d815-7164-449f-9afb-35f55a83bc06 · inbound

Meta-Learning Preferences for Multilingual LLM Alignment cites this paper.

Meta-Learning Preferences for Multilingual LLM Alignment Reinforcement Learning for LLM Post-Training: A Survey

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T05:36:16.410995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:36:16.410995Z digest=sha256:e6208c430ae61907bb24535518a9bc599371b76bece9fc76a5ded846e23390ba

Observation 59a1edaa-924c-4831-a5be-ff2c11c0bace · inbound

Sound Probabilistic Safety Bounds for Large Language Models cites this paper.

Sound Probabilistic Safety Bounds for Large Language Models Reinforcement Learning for LLM Post-Training: A Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T10:22:41.813908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:22:41.813908Z digest=sha256:66351f3c15b8e7caec49c3ee51ee4773f630ca057072d0d1f547ef2c87b43673

Observation d6cadd64-ae02-4da7-8972-71698e65181d · inbound

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment cites this paper.

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment Reinforcement Learning for LLM Post-Training: A Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-30T11:44:36.074744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T11:44:36.074744Z digest=sha256:6508cd441a93d503256e6c673422bbf8e70c134826af357540af8075e245abd5