Pith. sign in

Paper Citation Record · LEDGER

Natural Language Fine-Tuning

As of 11 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2412.20382.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20382 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:29:08.702803Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 765ec87f-85e9-4544-b8a2-d2a4c4a1cbbc · outbound

This paper cites Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future.

Natural Language Fine-Tuning Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:29:08.621228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:29:08.621228Z digest=sha256:e6e96260fac2c4b00ca56f18e3a8fd9365e56410301e62cb8dd9b33ff936ad8f

Observation f7d33d71-2741-4062-941e-e7cfbaaba0ef · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

Natural Language Fine-Tuning ReFT: Reasoning with Reinforced Fine-Tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:29:08.643231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:29:08.643231Z digest=sha256:6309f472ecf34f837f54947495fb9962cecbe13a8c815ddae54fe99b2f3ed8ac

Observation a7cba53f-2d86-49bb-b31a-247e4861d845 · outbound

This paper cites GPT-4 Technical Report.

Natural Language Fine-Tuning GPT-4 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:29:08.648900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:29:08.648900Z digest=sha256:a071acbe0811c5ab55d3572763f301aa41adaa757af1bd07bd375bde4a78f530

Observation d756caa0-f2fb-4bce-be55-08b9c331a589 · outbound

This paper cites Reinforcement learning from hu- man feedback research program,.

Natural Language Fine-Tuning Reinforcement learning from hu- man feedback research program,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:29:08.920202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:29:08.652900Z digest=sha256:128ae49cc89c91402dae02e957a7f687e92643256dbdf4996b1f616f86c91f6b

Observation cb228be1-5a26-440a-bedc-15678b933440 · outbound

This paper cites [Ouyang et al., 2022] Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al.

Natural Language Fine-Tuning [Ouyang et al., 2022] Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:29:08.907788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:29:08.657197Z digest=sha256:6a79e79c9a6c9573ae00fdbf39063dc022f4528c7618795e21c1f9b84e575d13

Observation 9cb83a03-c228-4c27-b273-504d872285f9 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Natural Language Fine-Tuning Code Llama: Open Foundation Models for Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:29:08.665935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:29:08.665935Z digest=sha256:a74616f1764b18c6e04b17abe8f5822054e0b4e7af5d863eabe1163b71713ebd

Observation 8cb1330c-674f-4e8c-be60-59a463d3754c · outbound

This paper cites Proximal Policy Optimization Algorithms.

Natural Language Fine-Tuning Proximal Policy Optimization Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:29:08.670100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:29:08.670100Z digest=sha256:8545475a9cedfd923575cb824ed8f76fddb87d747618ee22ee0203cf6f424ab3

Observation 43a126fe-1e80-49d4-80dd-0b96823585d8 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Natural Language Fine-Tuning Chain-of-thought prompting elicits reasoning in large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:29:08.682348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:29:08.682348Z digest=sha256:5d27f890992bfab5753821676f328bee1c59083d821a905292e517183de557e8

Observation 236528a0-cbca-49f0-9a94-92be0bb6102d · outbound

This paper cites Reasons to Reject? Aligning Language Models with Judgments.

Natural Language Fine-Tuning Reasons to Reject? Aligning Language Models with Judgments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:29:08.686241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:29:08.686241Z digest=sha256:f4a7ad34709ebb75cceb34055f8136604bad281b86a2d57e9453f8da21cbaa83

Observation e760f344-7ed6-400d-a704-02ceab086c57 · outbound

This paper cites Deep stable learning for out-of-distribution generalization.

Natural Language Fine-Tuning Deep stable learning for out-of-distribution generalization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:29:08.874066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:29:08.690351Z digest=sha256:1d0bcd949303b163e004b231e4f40dc1f0ce32420ed8155c1683b6f8c2edac4a

Observation 6b2af11e-e5ae-4fd5-a7ee-71fae1732d1c · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

Natural Language Fine-Tuning DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:29:08.694519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:29:08.694519Z digest=sha256:0aba9e2cac10e6f4a8ae71f56df96732e462e38b40016ed15d8a443f004fa4a7

Observation 73a0d80f-c045-4e1e-9971-1b50322df307 · outbound

This paper cites Instruction Suppose you are a math expert and you are presented with a math problem, a student’s response, and the correct answer.

Natural Language Fine-Tuning Instruction Suppose you are a math expert and you are presented with a math problem, a student’s response, and the correct answer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:29:08.860190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:29:08.698734Z digest=sha256:347febe1f7d79263d12979dd18f8f896c51c267eae71c3781ed56377c53f0c4f

Observation 06b2d7bd-eca1-4f4c-b314-024e629ef859 · outbound

This paper cites , 2020 ].

Natural Language Fine-Tuning , 2020 ]

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:29:08.845674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:29:08.702803Z digest=sha256:264e34619eb30f56929afe1e19a6f8da79e0472e5778efd9352fa901530497b5

Observation 51943ada-4e93-4881-abce-d2efda78f5f9 · outbound

This paper cites Trl: Transformer reinforcement learning.

Natural Language Fine-Tuning Trl: Transformer reinforcement learning

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:29:08.895108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:29:08.674197Z digest=sha256:88f010ec4e39e8f3083e0215a552d45593e5b4810334b4c18c9b335300f0c981

Observation 0fcc864f-840d-431f-89a3-d968d2426dcc · outbound

This paper cites Decoupled Weight Decay Regularization, January.

Natural Language Fine-Tuning Decoupled Weight Decay Regularization, January

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:29:08.932703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:29:08.639129Z digest=sha256:5b8ef1b373fcb3e75f3263d86900c46860f46c82f83ffc26cb63da016c7cdd2e

Observation e85af41d-ee57-4b22-b92b-5f428ec5cd10 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Natural Language Fine-Tuning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T23:29:08.678392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:29:08.678392Z digest=sha256:d744d96ddb67bc64d380e8e3522a9dd42161fa4b11f1681835fa730a9e94df55

Observation 80ee5de8-8d9f-444b-aad3-fd1f8015f637 · outbound

This paper cites Re- search on overfitting of deep learning.

Natural Language Fine-Tuning Re- search on overfitting of deep learning

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:29:08.944445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:29:08.634971Z digest=sha256:15a530215db22534861030042bff3ea806889c7a401f257d654ae5d07abaca06

Observation df902786-94fd-4181-93e2-7b54024678d0 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Natural Language Fine-Tuning From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T23:29:08.661148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:29:08.661148Z digest=sha256:eb974aa9053f79725fe4ec59ea1221bf150c31fb2b8e4a67d19ba259bca0a7fb

Observation fa261cbe-17d2-4642-8765-4b43140d3573 · outbound

This paper cites Training Verifiers to Solve Math Word Problems, November.

Natural Language Fine-Tuning Training Verifiers to Solve Math Word Problems, November

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:29:08.970995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:29:08.626224Z digest=sha256:87a4f3d168ea5029a8e6ff9fd7a8d316845e742bc92dfdec661278104a68982b

Observation 01886710-4719-4d2d-a968-1af17bedb062 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Natural Language Fine-Tuning Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:29:08.958156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:29:08.630732Z digest=sha256:9beb6ba59ab3058a00b6eaccf200ab62e6368d4253ba2cdeb8d04539d8a9f15b

Pith citing papers

No inbound Pith citation observations are available.