Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T13:17:42.848578Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 29 inbound Pith citation observations for arXiv:2307.04964.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T13:17:42.848578Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T08:15:59.314080Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T06:15:00.866473Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8c75b798-b3e1-4489-ab0a-e23faef0efde · outbound
Secrets of RLHF in Large Language Models Part I: PPO LLaMA: Open and Efficient Foundation Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c41d7993-64ca-4374-af59-1d3d6872d061 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 69dfe1fb-bb39-4474-a9a0-4fb6b7b47135 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Gpt-4 technical report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b3233dc6-fc14-453f-819c-bc31f858b810 · outbound
Secrets of RLHF in Large Language Models Part I: PPO A Survey of Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e5c94e01-3d74-488e-8ef2-b21f6025527d · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9ef53359-f728-4091-936a-4cd9ea33fbb8 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Instruction Tuning with GPT-4
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9420be41-fa2d-4017-8d71-c27338517d79 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Gulrajani, T
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 450942f6-a1f2-484d-86d8-91000a1df7a1 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 67047b39-235c-4e3a-b88e-505211a4a1c2 · outbound
Secrets of RLHF in Large Language Models Part I: PPO PaLM-E: An Embodied Multimodal Language Model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 063d1ef0-7e09-4872-a931-8d427161e3da · outbound
Secrets of RLHF in Large Language Models Part I: PPO Generative Agents: Interactive Simulacra of Human Behavior
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 90db619f-d530-4d5e-8a98-e01cadefdf5a · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8bbb306d-edac-4934-99a0-9bf0c388b194 · outbound
Secrets of RLHF in Large Language Models Part I: PPO LaMDA: Language Models for Dialog Applications
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 362e272c-ad1b-43b8-a6d9-1cff94be1aec · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 959517ac-717e-466e-92a3-f0789164ed41 · outbound
Secrets of RLHF in Large Language Models Part I: PPO On the Opportunities and Risks of Foundation Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1be1a84e-7f77-4bb1-920b-74cd8fd6b9ae · outbound
Secrets of RLHF in Large Language Models Part I: PPO Planning for agi and beyond
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3e0e8d1b-3579-4a52-87f2-c3192cdb82db · outbound
Secrets of RLHF in Large Language Models Part I: PPO Training language models to follow instructions with human feedback
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 88012d84-56eb-4107-816c-7c24261f9c3d · outbound
Secrets of RLHF in Large Language Models Part I: PPO Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e90edba7-e98e-4085-9bd4-d107eedad9a1 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Open-Chinese-LLaMA: Chinese large language model base generated through incremental pre-training on chinese datasets
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a52d831f-0e06-46f9-8460-32f56fa5f1b9 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 27c5086e-590a-45ed-b29b-b7a2c2194e99 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation abcd338c-bf78-4d39-aa02-6e24a222c534 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Belkada, K
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b8987e0e-67b4-4a9a-9d4a-8b52f2912a79 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a75b9212-8747-4e18-bc3d-0b6d587a30fe · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9ffa8ae5-82d4-4c32-a140-ccf029b722e0 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Fine-Tuning Language Models from Human Preferences
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c4918f68-6023-4133-9914-7d79d4998a27 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Ouyang, J
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 13b42ce3-24b3-4375-ad50-ad35d9adde35 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Kadavath, S
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e420e98f-deb5-4b75-b940-f55afb5930e4 · outbound
Secrets of RLHF in Large Language Models Part I: PPO A General Language Assistant as a Laboratory for Alignment
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4b09fc6d-a2d2-40ba-8eaf-036cd0064c13 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Raichuk, P
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bef6cfba-6e93-4065-b7e3-5e533fe3d3dd · outbound
Secrets of RLHF in Large Language Models Part I: PPO Ilyas, S
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 44eca8ee-1ff5-429a-aca5-9022f487e621 · outbound
Secrets of RLHF in Large Language Models Part I: PPO The Curious Case of Neural Text Degeneration
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0451a93a-93d7-4eba-853e-49fe3e44faa5 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b4ca9be8-c4da-4a10-8bf8-7f6f54702625 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dbb6f39c-977c-4500-a387-4a29d5386b26 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Levine, P
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dc6f1e59-ff0d-4fdb-a2b1-bc929e6d469b · outbound
Secrets of RLHF in Large Language Models Part I: PPO Wolski, P
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 87fc0409-5d9d-4f22-b811-80b23c0f40e8 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e0023d86-d8e4-4e5a-a8c3-f31381eacd33 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Kavukcuoglu, D
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 67435465-f740-45c3-a8b6-d6ae66cd97f0 · outbound
Secrets of RLHF in Large Language Models Part I: PPO J., Yiyuan Yang.Easy RL: Reinforcement Learning Tutorial
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6b7da4a7-22f3-4533-a316-20192625f5b9 · outbound
Secrets of RLHF in Large Language Models Part I: PPO McCann, L
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d0250bf3-b28c-4df3-bbea-c19bfef28078 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4674a8af-5b4c-41dd-bd60-2f696b8427a4 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Chiang, Y
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bef4b476-4310-4bb5-9d01-a8599bd2ddd6 · outbound
Secrets of RLHF in Large Language Models Part I: PPO The idea is that these organisms could have survived the journey through space and then established themselves on our planet
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fa60282d-836e-4e01-bd3a-50052dc98524 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Over time, these compounds would have organized themselves into more complex molecules, eventually leading to the formation of the first living cells
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3534006e-65a9-4099-b250-869414f4e831 · outbound
Secrets of RLHF in Large Language Models Part I: PPO These organisms were able to thrive in an environment devoid of sunlight, using chemical energy instead
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6f741e3f-c0db-416a-b8d4-47ccb5660883 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4e4b74ec-10a9-4637-8852-ca86eb6baa70 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f82109a0-35fe-4672-a8a3-ff0b5a9e003f · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3ec9ff71-fb82-4f7a-baa1-956c70c3d845 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4a0ffa7c-30a9-44c4-97c2-2f913f998ce0 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 035e814d-230c-49e8-a725-b013ab446ec8 · outbound
Secrets of RLHF in Large Language Models Part I: PPO This is just one example of many scams that prey on vulnerable older adults
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f3435e97-6136-4655-91ad-3be6fc2c52c3 · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 75ec9899-edcd-40e2-8855-d208388ede79 · outbound
Secrets of RLHF in Large Language Models Part I: PPO It can result in legal trouble, damage to your reputation, and harm to your relationships with others
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c4c5e115-b233-4864-9672-f126fbc4805a · outbound
Secrets of RLHF in Large Language Models Part I: PPO Taking advantage of their trust and naivety can be morally wrong
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7e6008d8-d889-45bf-993c-7f4a7593b6cc · outbound
Secrets of RLHF in Large Language Models Part I: PPO Table 2: Examples of models’ responses to the English prompt on helpful and harmful evaluation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 00ca29f3-1af6-4287-b6d6-bfb0e72eb7bf · outbound
Secrets of RLHF in Large Language Models Part I: PPO Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 41490a4e-bb89-416f-a14e-247a2c2e6e79 · outbound
Secrets of RLHF in Large Language Models Part I: PPO 非法使用信息网络,视情节及后果表现,处5日以上10日以 下的行政拘留;情节严重、后果严重的,处10日以上15日以下的行政拘留
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 671038ff-fc7d-4189-b76b-a81cfb8df0c6 · outbound
Secrets of RLHF in Large Language Models Part I: PPO second hand of love
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b5b79e4d-b95f-4f83-94a0-8fef0be7e85d · outbound
Secrets of RLHF in Large Language Models Part I: PPO Instead of rushing through life, take the time to savor the small things and appreciate the people around you
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 42d25b89-75e8-401d-a937-79bed76d39a0 · outbound
Secrets of RLHF in Large Language Models Part I: PPO When seen as a second hand of love, every moment counts, and we should make the most of them by being fully present and engaged in our relationships and experiences
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 67bede84-c1a4-47b7-985c-2d82f0f1bd4b · outbound
Secrets of RLHF in Large Language Models Part I: PPO We should focus on what truly matters to us and prioritize our time accordingly
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 06c76c60-4bd7-43d7-8570-9f03c4d87e0f · outbound
Secrets of RLHF in Large Language Models Part I: PPO The Wandering Earth
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a5fec650-f8f9-465a-b4eb-ff755aa6ef86 · inbound
InternLM2 Technical Report Secrets of RLHF in Large Language Models Part I: PPO
Reference 156
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c7d83599-890a-4bf7-885e-0fd052e2e2da · inbound
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Secrets of RLHF in Large Language Models Part I: PPO
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d8732b1e-acf2-448e-b000-e2ce46e13a6d · inbound
BalancedDPO: Adaptive Multi-Metric Alignment Secrets of RLHF in Large Language Models Part I: PPO
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aecf258c-2f26-4c3b-9268-43ef437e0746 · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Secrets of RLHF in Large Language Models Part I: PPO
Reference 185
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 079b57ab-66c1-488c-aa25-c81a6c23f8fc · inbound
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models Secrets of RLHF in Large Language Models Part I: PPO
Reference 115
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79a032b9-b362-4938-b476-819ec8889aea · inbound
Value Drifts: Tracing Value Alignment During LLM Post-Training Secrets of RLHF in Large Language Models Part I: PPO
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42533882-2aa0-4c01-8b88-00bab6aa5854 · inbound
Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models Secrets of RLHF in Large Language Models Part I: PPO
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a37ad172-a3aa-42b7-9bae-321b09a92d3d · inbound
Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models Secrets of RLHF in Large Language Models Part I: PPO
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 16af6309-c547-4d77-8f62-08b5bcb033f5 · inbound
Structure Matters: Evaluating Multi-Agents Orchestration in Generative Therapeutic Chatbots Secrets of RLHF in Large Language Models Part I: PPO
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6744e1cf-4792-4f3f-9a6b-98c5e803d947 · inbound
Joint Optimization of Multi-agent Memory System Secrets of RLHF in Large Language Models Part I: PPO
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fb0571e0-0a13-48c9-986c-0ce4a464119f · inbound
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning Secrets of RLHF in Large Language Models Part I: PPO
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0fb9b332-5e56-44d2-8074-21af8a30aee0 · inbound
Low-rank Optimization Trajectories Modeling for LLM RLVR Acceleration Secrets of RLHF in Large Language Models Part I: PPO
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7a6355db-9581-421a-af63-35552903dee1 · inbound
Representation-Guided Parameter-Efficient LLM Unlearning Secrets of RLHF in Large Language Models Part I: PPO
Reference 203
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 337a1bbe-7e7d-4c1c-99fd-a31b223f6a84 · inbound
EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training Secrets of RLHF in Large Language Models Part I: PPO
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dd53dbfc-81d9-4f7e-a992-c3e160139285 · inbound
Cost-Aware Learning Secrets of RLHF in Large Language Models Part I: PPO
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8ddb67de-8805-48a4-a6b6-7871e39b5c86 · inbound
Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning Secrets of RLHF in Large Language Models Part I: PPO
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bc8c70f2-f20d-4ef4-a46b-8675df69faad · inbound
WeatherSyn: An Instruction Tuning MLLM For Weather Forecasting Report Generation Secrets of RLHF in Large Language Models Part I: PPO
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 140bbb1e-2e88-4f88-9cd4-638c79a8512e · inbound
Reinforcement Learning for Scalable and Trustworthy Intelligent Systems Secrets of RLHF in Large Language Models Part I: PPO
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 637d5dbf-03c7-4652-a35d-d3da5fb642ba · inbound
Hint Tuning: Less Data Makes Better Reasoners Secrets of RLHF in Large Language Models Part I: PPO
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b199c1fa-6a92-43f7-a2c0-b59a76b8d585 · inbound
Hint Tuning: Less Data Makes Better Reasoners Secrets of RLHF in Large Language Models Part I: PPO
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 671a9382-df45-42b8-b36b-0bc2c338e1a5 · inbound
Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning Secrets of RLHF in Large Language Models Part I: PPO
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fa0af2ee-e070-4316-9378-9308c1a43bc9 · inbound
Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use Secrets of RLHF in Large Language Models Part I: PPO
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c6cc844f-c89b-4570-8f6c-5567af0040d1 · inbound
ASymPO: Asymmetric-Scale Policy Optimization for Asynchronous LLM Post-Training Without Behavior Information Secrets of RLHF in Large Language Models Part I: PPO
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c96ae485-9b02-4d5f-9a56-6fbc903912cd · inbound
DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity Secrets of RLHF in Large Language Models Part I: PPO
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 802fb54f-9835-46d3-9103-0cf7a54865d2 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Secrets of RLHF in Large Language Models Part I: PPO
Reference 274
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 87039985-0caf-465a-a8fa-279fcfc84daf · inbound
DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation Secrets of RLHF in Large Language Models Part I: PPO
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e97938e4-c485-40ed-901f-838cdbb0a09c · inbound
Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Secrets of RLHF in Large Language Models Part I: PPO
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6feb314c-d597-4fd9-8763-ced9bfa9ad1f · inbound
Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels Secrets of RLHF in Large Language Models Part I: PPO
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b45bfb3-0e3c-4249-b41b-dc6be31f009c · inbound
Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning Secrets of RLHF in Large Language Models Part I: PPO
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.