Pith. sign in

Paper Citation Record · LEDGER

Secrets of RLHF in Large Language Models Part I: PPO

As of 4 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 29 inbound Pith citation observations for arXiv:2307.04964.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.04964 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T13:17:42.848578Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T08:15:59.314080Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:15:00.866473Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact14
  • verified fuzzy27
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8c75b798-b3e1-4489-ab0a-e23faef0efde · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Secrets of RLHF in Large Language Models Part I: PPO LLaMA: Open and Efficient Foundation Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.968020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:6f557d0e173e90c72f76df85bcb7421a4c38e5ef39ed4a5e2b70eb5b3a24ef6c

Observation c41d7993-64ca-4374-af59-1d3d6872d061 · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.093995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:275140b97aa3185da1a54644a2b0cdcd303ff59f8bb40d849bca18f6d60c9848

Observation 69dfe1fb-bb39-4474-a9a0-4fb6b7b47135 · outbound

This paper cites Gpt-4 technical report.

Secrets of RLHF in Large Language Models Part I: PPO Gpt-4 technical report

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.096993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:335c4477c0ff2cf14c695eb86118497a9bab4634efd1b65f5d90c8cd8c86244b

Observation b3233dc6-fc14-453f-819c-bc31f858b810 · outbound

This paper cites A Survey of Large Language Models.

Secrets of RLHF in Large Language Models Part I: PPO A Survey of Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.893700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:91be880c2b7a5fc620212f90558df6c1433465d300d5c079cc6a7cdd612b4f2a

Observation e5c94e01-3d74-488e-8ef2-b21f6025527d · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.100456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:00f464eb7c2c32e561b6e03a113c409f2747271f284584cda180610b9342fa93

Observation 9ef53359-f728-4091-936a-4cd9ea33fbb8 · outbound

This paper cites Instruction Tuning with GPT-4.

Secrets of RLHF in Large Language Models Part I: PPO Instruction Tuning with GPT-4

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.956775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:35d00a5e89f741329612e5c0c6a17e4684a7eda478dd87d4500bed6cd810ac49

Observation 9420be41-fa2d-4017-8d71-c27338517d79 · outbound

This paper cites Gulrajani, T.

Secrets of RLHF in Large Language Models Part I: PPO Gulrajani, T

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.107279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:1f99e92715bc6302c239600e9c515900db5502ade9405f6ee666eaeb6ad8e650

Observation 450942f6-a1f2-484d-86d8-91000a1df7a1 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Secrets of RLHF in Large Language Models Part I: PPO Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.913043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:63c1c712b60bf19677ecfb0afc263d4b60207e4e01265ee4f6b0366c7ae7532f

Observation 67047b39-235c-4e3a-b88e-505211a4a1c2 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Secrets of RLHF in Large Language Models Part I: PPO PaLM-E: An Embodied Multimodal Language Model

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.937746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:be1d4dd92166e3af834346a359c878b531e430ddd350f46100798da12acc00f4

Observation 063d1ef0-7e09-4872-a931-8d427161e3da · outbound

This paper cites Generative Agents: Interactive Simulacra of Human Behavior.

Secrets of RLHF in Large Language Models Part I: PPO Generative Agents: Interactive Simulacra of Human Behavior

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.947701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:8b4dc25661879a2bf4b5be4856314143673a3f934f95ebef42f50c6b7ed2f583

Observation 90db619f-d530-4d5e-8a98-e01cadefdf5a · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.110813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:469188ec225adedb5185b6e477f0ce071ccfcd0adc34c0f8723fdae12131277d

Observation 8bbb306d-edac-4934-99a0-9bf0c388b194 · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

Secrets of RLHF in Large Language Models Part I: PPO LaMDA: Language Models for Dialog Applications

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.887801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:9124508d7a1487d3751e1f93e2245670267414df42c92f1d3141a1b3ea355559

Observation 362e272c-ad1b-43b8-a6d9-1cff94be1aec · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.113927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:17ec02e9c57792f6309d58f070b2517edda6e16e514164dc5937a29c1d352283

Observation 959517ac-717e-466e-92a3-f0789164ed41 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Secrets of RLHF in Large Language Models Part I: PPO On the Opportunities and Risks of Foundation Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.899373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:bcde89ba1f361c14da8d44246a0684cdd5a63dcba43c1a9563a9395f11b14e50

Observation 1be1a84e-7f77-4bb1-920b-74cd8fd6b9ae · outbound

This paper cites Planning for agi and beyond.

Secrets of RLHF in Large Language Models Part I: PPO Planning for agi and beyond

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.117595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:3f3a21387da12855b5c718d2ffd84a4c43bb7177c14d97df093ec98603c2e708

Observation 3e0e8d1b-3579-4a52-87f2-c3192cdb82db · outbound

This paper cites Training language models to follow instructions with human feedback.

Secrets of RLHF in Large Language Models Part I: PPO Training language models to follow instructions with human feedback

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.920431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:35ef562313493ac61f785c8bded8d79a40fe13032ce49f7521ba57d10e954598

Observation 88012d84-56eb-4107-816c-7c24261f9c3d · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Secrets of RLHF in Large Language Models Part I: PPO Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.931062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:c61d9b3754bf7cac700cbd22ef56823b8dc2c482240cfd19de67d59884508d42

Observation e90edba7-e98e-4085-9bd4-d107eedad9a1 · outbound

This paper cites Open-Chinese-LLaMA: Chinese large language model base generated through incremental pre-training on chinese datasets.

Secrets of RLHF in Large Language Models Part I: PPO Open-Chinese-LLaMA: Chinese large language model base generated through incremental pre-training on chinese datasets

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.121043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:31871306704d5b734673c5c186ac53a569ab88ed6e2084f52907c70a69533889

Observation a52d831f-0e06-46f9-8460-32f56fa5f1b9 · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.123770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:1878923ed7468a320c63993b345f52b9ffb5159c1494ce637db203f36dd3809f

Observation 27c5086e-590a-45ed-b29b-b7a2c2194e99 · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.126563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:691aa4e1be58fe398536305cd5d2401bcdb1d47f84c055e8a1265a223f5b4b0b

Observation abcd338c-bf78-4d39-aa02-6e24a222c534 · outbound

This paper cites Belkada, K.

Secrets of RLHF in Large Language Models Part I: PPO Belkada, K

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.130019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:05a4c5971c4785fa7f0c9b7e2a0fb41b466c2579102306d9e2d06dea6db8eee7

Observation b8987e0e-67b4-4a9a-9d4a-8b52f2912a79 · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.133088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:790e04e33dc5de01b6c5a5dc0dff3022529cf5eb36f64508fe800b25756273f0

Observation a75b9212-8747-4e18-bc3d-0b6d587a30fe · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.135888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:49f7eb6a024b7d30a703b7e99eba517034773c012a22708b602369e41e2f9d58

Observation 9ffa8ae5-82d4-4c32-a140-ccf029b722e0 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Secrets of RLHF in Large Language Models Part I: PPO Fine-Tuning Language Models from Human Preferences

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.904891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:79c91a28fab4976679100288cd9163e6f732b3780dc4f169ed08792372e8b9dc

Observation c4918f68-6023-4133-9914-7d79d4998a27 · outbound

This paper cites Ouyang, J.

Secrets of RLHF in Large Language Models Part I: PPO Ouyang, J

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:42.970881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:795d6485c089160c4c95b64dd48bca49934dae56be3668ceb031d3f01fbd1194

Observation 13b42ce3-24b3-4375-ad50-ad35d9adde35 · outbound

This paper cites Kadavath, S.

Secrets of RLHF in Large Language Models Part I: PPO Kadavath, S

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:42.973536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:96c680fd49ceed826a6d5a0f9eb614f1b6936c6fb9b0c497dd60b8719efdee64

Observation e420e98f-deb5-4b75-b940-f55afb5930e4 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Secrets of RLHF in Large Language Models Part I: PPO A General Language Assistant as a Laboratory for Alignment

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.926079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:cceea2f2088c7aad40a9abdb032575cde784e4fc48bbb853d7dd6ea2614cb9bd

Observation 4b09fc6d-a2d2-40ba-8eaf-036cd0064c13 · outbound

This paper cites Raichuk, P.

Secrets of RLHF in Large Language Models Part I: PPO Raichuk, P

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:42.976099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:231ad92ffd17680c03c8138a1fb89aa1f93ea53499b9b3b70422d0c08c4a508d

Observation bef6cfba-6e93-4065-b7e3-5e533fe3d3dd · outbound

This paper cites Ilyas, S.

Secrets of RLHF in Large Language Models Part I: PPO Ilyas, S

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:42.978598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:c59d6bdc75acda39b8a1bb13327addce3fddbae746c0823a63060841748df653

Observation 44eca8ee-1ff5-429a-aca5-9022f487e621 · outbound

This paper cites The Curious Case of Neural Text Degeneration.

Secrets of RLHF in Large Language Models Part I: PPO The Curious Case of Neural Text Degeneration

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.942850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:16462a7cfcca68c2e9ae4ae1a2d4e70d4548b228e57b120460952c83a0283dad

Observation 0451a93a-93d7-4eba-853e-49fe3e44faa5 · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:42.981207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:8d97ae77b51764b2a0bd160e9b5dd24d5f3ae818b7b5c65881c49bd6f3463b7b

Observation b4ca9be8-c4da-4a10-8bf8-7f6f54702625 · outbound

This paper cites Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog.

Secrets of RLHF in Large Language Models Part I: PPO Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:17:42.952535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:ceb7174f524cf297534c769f441163b5385340e3534bcaab1345fe5e9661dbb6

Observation dbb6f39c-977c-4500-a387-4a29d5386b26 · outbound

This paper cites Levine, P.

Secrets of RLHF in Large Language Models Part I: PPO Levine, P

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:42.983742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:47a83d5964ae3ce1f273b9d1ec43b3802e33fabb895292604f27d6116ea938e7

Observation dc6f1e59-ff0d-4fdb-a2b1-bc929e6d469b · outbound

This paper cites Wolski, P.

Secrets of RLHF in Large Language Models Part I: PPO Wolski, P

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:42.986362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:e5e832df501fc909f50a776e371d6e50e05fdc8fd463a7bcc6cf271c63076e6c

Observation 87fc0409-5d9d-4f22-b811-80b23c0f40e8 · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:42.989057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:e65f4324dd1bae3c2a90230d21b21bb969a91eba499d7dfd819febc95c003ba4

Observation e0023d86-d8e4-4e5a-a8c3-f31381eacd33 · outbound

This paper cites Kavukcuoglu, D.

Secrets of RLHF in Large Language Models Part I: PPO Kavukcuoglu, D

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:42.991580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:c61af1b8f77d9d27513643d2863b562d6b329129011757c7320a5f851912a38d

Observation 67435465-f740-45c3-a8b6-d6ae66cd97f0 · outbound

This paper cites J., Yiyuan Yang.Easy RL: Reinforcement Learning Tutorial.

Secrets of RLHF in Large Language Models Part I: PPO J., Yiyuan Yang.Easy RL: Reinforcement Learning Tutorial

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:42.994445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:c14e567b424f4d61cec8293416489eef89614fb069c2eb8a01714b14013e068f

Observation 6b7da4a7-22f3-4533-a316-20192625f5b9 · outbound

This paper cites McCann, L.

Secrets of RLHF in Large Language Models Part I: PPO McCann, L

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:42.998877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:cbfe102e2ac7570acb8400a26c6f88f469a8a8d8ef287017f7066769e1690e5f

Observation d0250bf3-b28c-4df3-bbea-c19bfef28078 · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.001939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:80dfc74b0f7b8a5c342701e5cc6f9a060c9789f034b3059a6f2c4bdbb24724e7

Observation 4674a8af-5b4c-41dd-bd60-2f696b8427a4 · outbound

This paper cites Chiang, Y.

Secrets of RLHF in Large Language Models Part I: PPO Chiang, Y

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.010143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:603815e2d946b49f4a7ddb2920c5300fe4a6d53fe348a7728ba5901c393ca3dd

Observation bef4b476-4310-4bb5-9d01-a8599bd2ddd6 · outbound

This paper cites The idea is that these organisms could have survived the journey through space and then established themselves on our planet.

Secrets of RLHF in Large Language Models Part I: PPO The idea is that these organisms could have survived the journey through space and then established themselves on our planet

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.014840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:dd635e4fc3661cfd5e1328b7a6c6b59db52a25f5f2a3113a1d034a5babbe3f02

Observation fa60282d-836e-4e01-bd3a-50052dc98524 · outbound

This paper cites Over time, these compounds would have organized themselves into more complex molecules, eventually leading to the formation of the first living cells.

Secrets of RLHF in Large Language Models Part I: PPO Over time, these compounds would have organized themselves into more complex molecules, eventually leading to the formation of the first living cells

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.021831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:be8d2c0a4982af8d41da8e9784c25c0c347d0ac50c72257849d03e6a976465f6

Observation 3534006e-65a9-4099-b250-869414f4e831 · outbound

This paper cites These organisms were able to thrive in an environment devoid of sunlight, using chemical energy instead.

Secrets of RLHF in Large Language Models Part I: PPO These organisms were able to thrive in an environment devoid of sunlight, using chemical energy instead

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.025162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:390ca7ba35549cc188767d41950a80c419382bed2447edf12520f44b24e84008

Observation 6f741e3f-c0db-416a-b8d4-47ccb5660883 · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.038985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:8fe0d00cdc507d7d15b8b9f2e80ee3e04172fde7a14aa5d53dc9824a77d04791

Observation 4e4b74ec-10a9-4637-8852-ca86eb6baa70 · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.042349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:5b78c73b99761d43b8b82dba2872770a27eb34de28a3f8132b54caf97fa8dedd

Observation f82109a0-35fe-4672-a8a3-ff0b5a9e003f · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.045564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:3c0ab08ee7ad64b118a59793b3e2ec8c183fd689cfd8f37504a9aea6f32d115a

Observation 3ec9ff71-fb82-4f7a-baa1-956c70c3d845 · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.048597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:28f350ed7d6af5937c3c3ad0e3c2743082336e0d78499d6ad77695c0e9d39ea7

Observation 4a0ffa7c-30a9-44c4-97c2-2f913f998ce0 · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.051474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:90f1340f03a5387d4b573e461ce4642f6c941b2d883ab2db2b1ff622338b2098

Observation 035e814d-230c-49e8-a725-b013ab446ec8 · outbound

This paper cites This is just one example of many scams that prey on vulnerable older adults.

Secrets of RLHF in Large Language Models Part I: PPO This is just one example of many scams that prey on vulnerable older adults

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.054819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:76bc22a5a7324931fa06357ef946c9195fb5a39500e09fb5ab7ce44f99e04a27

Observation f3435e97-6136-4655-91ad-3be6fc2c52c3 · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.057937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:d886dc9bee5788d7df885e3038a02105edded9f0099a6b52a5cb203525b477fd

Observation 75ec9899-edcd-40e2-8855-d208388ede79 · outbound

This paper cites It can result in legal trouble, damage to your reputation, and harm to your relationships with others.

Secrets of RLHF in Large Language Models Part I: PPO It can result in legal trouble, damage to your reputation, and harm to your relationships with others

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.060918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:cd710725864ee6dea26a64ffb9a3f72dc0866ce76160460f13b757bc4f0b4c8b

Observation c4c5e115-b233-4864-9672-f126fbc4805a · outbound

This paper cites Taking advantage of their trust and naivety can be morally wrong.

Secrets of RLHF in Large Language Models Part I: PPO Taking advantage of their trust and naivety can be morally wrong

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.063957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:12330ce020ec15a853700e0696124b1269268fd3a062a13da21eac28326d10a4

Observation 7e6008d8-d889-45bf-993c-7f4a7593b6cc · outbound

This paper cites Table 2: Examples of models’ responses to the English prompt on helpful and harmful evaluation.

Secrets of RLHF in Large Language Models Part I: PPO Table 2: Examples of models’ responses to the English prompt on helpful and harmful evaluation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.067928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:e8d718392d7708703519be725049a60666ad228a0568472105cc902538c24ee7

Observation 00ca29f3-1af6-4287-b6d6-bfb0e72eb7bf · outbound

This paper cites an unresolved cited work.

Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-17T13:17:43.071241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:02be10cc8a23384c98e0ba0182817facae9784fc13bf12daa812ca518ebaa811

Observation 41490a4e-bb89-416f-a14e-247a2c2e6e79 · outbound

This paper cites 非法使用信息网络,视情节及后果表现,处5日以上10日以 下的行政拘留;情节严重、后果严重的,处10日以上15日以下的行政拘留.

Secrets of RLHF in Large Language Models Part I: PPO 非法使用信息网络,视情节及后果表现,处5日以上10日以 下的行政拘留;情节严重、后果严重的,处10日以上15日以下的行政拘留

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.074603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:43863198aba4258b68a1c412b46d49a7c81b35865a24842d822faaf5355af1bf

Observation 671038ff-fc7d-4189-b76b-a81cfb8df0c6 · outbound

This paper cites second hand of love.

Secrets of RLHF in Large Language Models Part I: PPO second hand of love

Reference 56

Resolution
malformed identifier
raw_fallback, observed 2026-05-17T13:17:43.079136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:bdba33d2b0c85f2dd6136cbaaa977c74699deeb7953969ec0fd0527b00f1a165

Observation b5b79e4d-b95f-4f83-94a0-8fef0be7e85d · outbound

This paper cites Instead of rushing through life, take the time to savor the small things and appreciate the people around you.

Secrets of RLHF in Large Language Models Part I: PPO Instead of rushing through life, take the time to savor the small things and appreciate the people around you

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.082149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:7661deba2368c0ec9539f419cd67c79252bc68fdc04c61924c07fa43802b0393

Observation 42d25b89-75e8-401d-a937-79bed76d39a0 · outbound

This paper cites When seen as a second hand of love, every moment counts, and we should make the most of them by being fully present and engaged in our relationships and experiences.

Secrets of RLHF in Large Language Models Part I: PPO When seen as a second hand of love, every moment counts, and we should make the most of them by being fully present and engaged in our relationships and experiences

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.085204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:10263902024321d970264021c76d23d0a0f424d69ae548f7a6d65f2409e3ca1b

Observation 67bede84-c1a4-47b7-985c-2d82f0f1bd4b · outbound

This paper cites We should focus on what truly matters to us and prioritize our time accordingly.

Secrets of RLHF in Large Language Models Part I: PPO We should focus on what truly matters to us and prioritize our time accordingly

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.088537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:392a754ee122ea172b6a1415fb4a1ac036f68da22aa7eab21f1944f02600ce24

Observation 06c76c60-4bd7-43d7-8570-9f03c4d87e0f · outbound

This paper cites The Wandering Earth.

Secrets of RLHF in Large Language Models Part I: PPO The Wandering Earth

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:17:43.091413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-17T13:17:42.848578Z digest=sha256:114adc24166f0eeb384102f3d62d469c0d9917159284515642d39b6a66b653e3

Pith citing papers

Observation a5fec650-f8f9-465a-b4eb-ff755aa6ef86 · inbound

InternLM2 Technical Report cites this paper.

InternLM2 Technical Report Secrets of RLHF in Large Language Models Part I: PPO

Reference 156

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-15T11:44:38.066501Z digest=sha256:6434cd84bb90bff145b709a615b0b405d6bddc9a98ec56af8e0e4a8086790feb

Observation c7d83599-890a-4bf7-885e-0fd052e2e2da · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Secrets of RLHF in Large Language Models Part I: PPO

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:862ed35cba76f299bb45d66a86661abcaf907cb8d5631d2a162a93f3615cb0c3

Observation d8732b1e-acf2-448e-b000-e2ce46e13a6d · inbound

BalancedDPO: Adaptive Multi-Metric Alignment cites this paper.

BalancedDPO: Adaptive Multi-Metric Alignment Secrets of RLHF in Large Language Models Part I: PPO

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:42:16.511659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T23:37:55.154902Z digest=sha256:f6bee4239eb94416866678520d18e5bef349ebdf48d060d27fdc043efe27ac7c

Observation aecf258c-2f26-4c3b-9268-43ef437e0746 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Secrets of RLHF in Large Language Models Part I: PPO

Reference 185

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:fe1facd60a09a10e4226f0c074bfa65f9677fe62f720e5b02f0de47e4e6434b6

Observation 079b57ab-66c1-488c-aa25-c81a6c23f8fc · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models Secrets of RLHF in Large Language Models Part I: PPO

Reference 115

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:59.314080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:59.314080Z digest=sha256:a923c08f762dce90015aacdb0a2e2bd2d9bdeffb6f2e55224b4c555c4d8de9a9

Observation 79a032b9-b362-4938-b476-819ec8889aea · inbound

Value Drifts: Tracing Value Alignment During LLM Post-Training cites this paper.

Value Drifts: Tracing Value Alignment During LLM Post-Training Secrets of RLHF in Large Language Models Part I: PPO

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:37.066762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:37.066762Z digest=sha256:2ddc9bd6ecdd6cc1a126f9994e60a80479dfa5a3b6ef5855442da88c4d583741

Observation 42533882-2aa0-4c01-8b88-00bab6aa5854 · inbound

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models cites this paper.

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models Secrets of RLHF in Large Language Models Part I: PPO

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T16:25:03.743594Z digest=sha256:0b307e357b0066f4be1a3cc20918fe63797bf9118c7790200ef1f1186f9f815a

Observation a37ad172-a3aa-42b7-9bae-321b09a92d3d · inbound

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models cites this paper.

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models Secrets of RLHF in Large Language Models Part I: PPO

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-05-21T16:04:14.715694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T16:01:52.150950Z digest=sha256:f001d1e0fe4df91a5bd659342dc0e4e02eda7e386e0c09a5e64137b37cd05df5

Observation 16af6309-c547-4d77-8f62-08b5bcb033f5 · inbound

Structure Matters: Evaluating Multi-Agents Orchestration in Generative Therapeutic Chatbots cites this paper.

Structure Matters: Evaluating Multi-Agents Orchestration in Generative Therapeutic Chatbots Secrets of RLHF in Large Language Models Part I: PPO

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T17:59:40.966701Z digest=sha256:e537dc1f1c4083562470bffb6fd758e4c6d1d5f75022637a6f98c8483a98111f

Observation 6744e1cf-4792-4f3f-9a6b-98c5e803d947 · inbound

Joint Optimization of Multi-agent Memory System cites this paper.

Joint Optimization of Multi-agent Memory System Secrets of RLHF in Large Language Models Part I: PPO

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T12:12:35.056095Z digest=sha256:af55a3b6b29ea27f02f49a737e737275b7303d0dae781501a5ca8c47fe635df9

Observation fb0571e0-0a13-48c9-986c-0ce4a464119f · inbound

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning cites this paper.

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning Secrets of RLHF in Large Language Models Part I: PPO

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:28:58.515666Z digest=sha256:46b36aabcfde71619eb869c93091f866bb905603aed32efe9a327d64044945ef

Observation 0fb9b332-5e56-44d2-8074-21af8a30aee0 · inbound

Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration cites this paper.

Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration Secrets of RLHF in Large Language Models Part I: PPO

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:07:20.180349Z digest=sha256:978edacf8f2cab84125478d6219d4b2bef98b12fe6cd2c6d5169bc03f45153f2

Observation 7a6355db-9581-421a-af63-35552903dee1 · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning Secrets of RLHF in Large Language Models Part I: PPO

Reference 203

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:2b7e8f2d5b47bc6130f96792eb44e4b4aa1fa08c9d571c6d062e9d4c65a67a8d

Observation 337a1bbe-7e7d-4c1c-99fd-a31b223f6a84 · inbound

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training cites this paper.

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Secrets of RLHF in Large Language Models Part I: PPO

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T03:00:10.404757Z digest=sha256:0af78f91ec0e231d99e3195a6a167d0d066bc07eef399d03b2011a3ae1d67d2d

Observation dd53dbfc-81d9-4f7e-a992-c3e160139285 · inbound

Cost-Aware Learning cites this paper.

Cost-Aware Learning Secrets of RLHF in Large Language Models Part I: PPO

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-07T05:11:01.131590Z digest=sha256:5ac2ab6d9667e1f64948af0cd25c473a2d99596e1194c8904ce462b4ec06bd56

Observation 8ddb67de-8805-48a4-a6b6-7871e39b5c86 · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning Secrets of RLHF in Large Language Models Part I: PPO

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:f9115e8a9e42dbcbf404411ac945482a92f232a36f0bbac2da7c41f03a9ae105

Observation bc8c70f2-f20d-4ef4-a46b-8675df69faad · inbound

WeatherSyn: An Instruction Tuning MLLM For Weather Forecasting Report Generation cites this paper.

WeatherSyn: An Instruction Tuning MLLM For Weather Forecasting Report Generation Secrets of RLHF in Large Language Models Part I: PPO

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-11T01:55:07.218958Z digest=sha256:30097cb35773b6fdc33a140cfd6a18ee86b0c0b6a782dc45a905a6636d9f99d9

Observation 140bbb1e-2e88-4f88-9cd4-638c79a8512e · inbound

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems cites this paper.

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems Secrets of RLHF in Large Language Models Part I: PPO

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:47:40.772146Z digest=sha256:67db0d89e3389d4be431e33630746deec02ade57af6ed2d5b3d33148ce8d5217

Observation 637d5dbf-03c7-4652-a35d-d3da5fb642ba · inbound

Hint Tuning: Less Data Makes Better Reasoners cites this paper.

Hint Tuning: Less Data Makes Better Reasoners Secrets of RLHF in Large Language Models Part I: PPO

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T00:54:13.146373Z digest=sha256:23d7374fccf5fa299978b8582c35aac3a0460ec1539de0443187c509060137e7

Observation b199c1fa-6a92-43f7-a2c0-b59a76b8d585 · inbound

Hint Tuning: Less Data Makes Better Reasoners cites this paper.

Hint Tuning: Less Data Makes Better Reasoners Secrets of RLHF in Large Language Models Part I: PPO

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:35:07.069222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T23:34:43.785312Z digest=sha256:5a5ccd36c3fc596e1f6041036c887aa907c6a02e35166c01c5bdcdab7785e3f2

Observation 671a9382-df45-42b8-b36b-0bc2c338e1a5 · inbound

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning cites this paper.

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning Secrets of RLHF in Large Language Models Part I: PPO

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:17:43.137115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T02:19:27.345348Z digest=sha256:0739ef27c96816b364d92d05748f4e3ca834ce11ec67ec052eb72958ebb867ea

Observation fa0af2ee-e070-4316-9378-9308c1a43bc9 · inbound

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use cites this paper.

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use Secrets of RLHF in Large Language Models Part I: PPO

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:53:28.605328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T13:49:58.677299Z digest=sha256:d7729114a2d2aea07ec773196b84c60b9936376df98e45d58619e67cab1430d7

Observation c6cc844f-c89b-4570-8f6c-5567af0040d1 · inbound

ASymPO: Asymmetric-Scale Policy Optimization for Asynchronous LLM Post-Training Without Behavior Information cites this paper.

ASymPO: Asymmetric-Scale Policy Optimization for Asynchronous LLM Post-Training Without Behavior Information Secrets of RLHF in Large Language Models Part I: PPO

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T02:06:27.471531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T11:11:41.433882Z digest=sha256:7dcc07df172c6f2a267c87b18875a74753cc41cd2d5e5bea3e5dc07e93fae33a

Observation c96ae485-9b02-4d5f-9a56-6fbc903912cd · inbound

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity cites this paper.

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity Secrets of RLHF in Large Language Models Part I: PPO

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-27T17:31:06.955538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T17:26:17.072017Z digest=sha256:1199e3c74a665a9d38b62a15e0f0fbca6147a783acfbdacdf5bcb89a9078d778

Observation 802fb54f-9835-46d3-9103-0cf7a54865d2 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Secrets of RLHF in Large Language Models Part I: PPO

Reference 274

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:09:40.639382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:73b749358dafb9d4e51d48479f1f2b4c34a195b40b7261e282f0e542e0161d75

Observation 87039985-0caf-465a-a8fa-279fcfc84daf · inbound

DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation cites this paper.

DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation Secrets of RLHF in Large Language Models Part I: PPO

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:39:45.593093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T08:37:11.052296Z digest=sha256:b72a65eb06636393d426f6885e92944ed042ea1a74911dbb9aa970189b7be840

Observation e97938e4-c485-40ed-901f-838cdbb0a09c · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Secrets of RLHF in Large Language Models Part I: PPO

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T15:37:01.391649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:37:01.391649Z digest=sha256:90cc001477d50367eccef4c890ed713629a3973e83c3e14297ab1f2cd0f4cc52

Observation 6feb314c-d597-4fd9-8763-ced9bfa9ad1f · inbound

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels cites this paper.

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Secrets of RLHF in Large Language Models Part I: PPO

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T08:00:53.466595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:00:53.466595Z digest=sha256:9febd5ce9b77fb956882c2ebfecec95d00ec246685ae5afa345659344bdcf2b5

Observation 5b45bfb3-0e3c-4249-b41b-dc6be31f009c · inbound

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning cites this paper.

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning Secrets of RLHF in Large Language Models Part I: PPO

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-31T01:31:10.167491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T01:31:10.167491Z digest=sha256:8142450aca0d0537213c3e3dd92d3288256008fa9dad7e84ac560443ff769d95