Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:29:08.702803Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2412.20382.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:29:08.702803Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 765ec87f-85e9-4544-b8a2-d2a4c4a1cbbc · outbound
Natural Language Fine-Tuning Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7d33d71-2741-4062-941e-e7cfbaaba0ef · outbound
Natural Language Fine-Tuning ReFT: Reasoning with Reinforced Fine-Tuning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7cba53f-2d86-49bb-b31a-247e4861d845 · outbound
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d756caa0-f2fb-4bce-be55-08b9c331a589 · outbound
Natural Language Fine-Tuning Reinforcement learning from hu- man feedback research program,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cb228be1-5a26-440a-bedc-15678b933440 · outbound
Natural Language Fine-Tuning [Ouyang et al., 2022] Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9cb83a03-c228-4c27-b273-504d872285f9 · outbound
Natural Language Fine-Tuning Code Llama: Open Foundation Models for Code
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cb1330c-674f-4e8c-be60-59a463d3754c · outbound
Natural Language Fine-Tuning Proximal Policy Optimization Algorithms
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43a126fe-1e80-49d4-80dd-0b96823585d8 · outbound
Natural Language Fine-Tuning Chain-of-thought prompting elicits reasoning in large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 236528a0-cbca-49f0-9a94-92be0bb6102d · outbound
Natural Language Fine-Tuning Reasons to Reject? Aligning Language Models with Judgments
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e760f344-7ed6-400d-a704-02ceab086c57 · outbound
Natural Language Fine-Tuning Deep stable learning for out-of-distribution generalization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6b2af11e-e5ae-4fd5-a7ee-71fae1732d1c · outbound
Natural Language Fine-Tuning DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73a0d80f-c045-4e1e-9971-1b50322df307 · outbound
Natural Language Fine-Tuning Instruction Suppose you are a math expert and you are presented with a math problem, a student’s response, and the correct answer
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 06b2d7bd-eca1-4f4c-b314-024e629ef859 · outbound
Natural Language Fine-Tuning , 2020 ]
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 51943ada-4e93-4881-abce-d2efda78f5f9 · outbound
Natural Language Fine-Tuning Trl: Transformer reinforcement learning
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0fcc864f-840d-431f-89a3-d968d2426dcc · outbound
Natural Language Fine-Tuning Decoupled Weight Decay Regularization, January
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e85af41d-ee57-4b22-b92b-5f428ec5cd10 · outbound
Natural Language Fine-Tuning Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80ee5de8-8d9f-444b-aad3-fd1f8015f637 · outbound
Natural Language Fine-Tuning Re- search on overfitting of deep learning
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation df902786-94fd-4181-93e2-7b54024678d0 · outbound
Natural Language Fine-Tuning From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa261cbe-17d2-4642-8765-4b43140d3573 · outbound
Natural Language Fine-Tuning Training Verifiers to Solve Math Word Problems, November
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 01886710-4719-4d2d-a968-1af17bedb062 · outbound
Natural Language Fine-Tuning Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
No inbound Pith citation observations are available.