Pith. sign in

Paper Citation Record · LEDGER

Reinforcement learning fine-tuning of language model for instruction following and math reasoning

As of 8 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2506.21560.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21560 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:37:30.520795Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 67b22cb6-de3c-4a24-b323-b1a84269d16e · outbound

This paper cites arXiv preprint arXiv:2412.15287 (2024).

Reinforcement learning fine-tuning of language model for instruction following and math reasoning arXiv preprint arXiv:2412.15287 (2024)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.512978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.512978Z digest=sha256:528328c95bfbccc4b67cefc58e2c9d0c4cc74f363f83a6ec624ca6222dfe5ef5

Observation b4301b55-41a7-4964-a3b9-40574ff41443 · outbound

This paper cites Stream of Search (SoS): Learning to Search in Language.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Stream of Search (SoS): Learning to Search in Language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.746418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.746418Z digest=sha256:75d7b9614c5dedb94b1cb8543a986e5ddd8f67be2a06ca865fb697bb094cdd80

Observation 3b9a162a-60ac-446a-8ddc-30c8f0d508da · outbound

This paper cites RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.805270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.805270Z digest=sha256:16431a3a8b859fad269227fdee9d81780e8bc5216f0a2c86f794b1905bdf9af5

Observation 46b454a4-95ff-474f-bee4-d78c48d97283 · outbound

This paper cites Qwen2.5-Coder Technical Report.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Qwen2.5-Coder Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.960575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.960575Z digest=sha256:a40ea8e87cedbec0e3c5905290ae3db8c720eb243287c92901c0500f3d1cc351

Observation 59233e2a-8955-4a6b-a21e-f05e28ca516c · outbound

This paper cites GPT-4o System Card.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning GPT-4o System Card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:30.042682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:30.042682Z digest=sha256:8223e3fd9e08a6cb48f6a92523d8641c1ca1349155104bd3a67225154f8fbfda

Observation 9e941f99-4fb5-4202-af35-0900f6a5caa0 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Self-Refine: Iterative Refinement with Self-Feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:30.178554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:30.178554Z digest=sha256:194459104b1e1d29d2d138402982f150a776eac62bf6f906f8a1e31179368b74

Observation 259ad9ec-5d56-40a1-b865-db09449d6c89 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:30.234365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:30.234365Z digest=sha256:51dfb3cdb21214dac41a2f63f91a3fa6961be17404847974f10a64387dfa016c

Observation 1ac134b5-0f06-4d92-bcf6-248a6fe79e99 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:30.328965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:30.328965Z digest=sha256:1bb990531d973be8169eb59078c50eb68fe5ad737bdbe67d92b1dda77f480af4

Observation a2596d1b-9eaa-4455-a10c-3bd2053ee55f · outbound

This paper cites Advances in Neural Information Processing Systems 36 (2023), 68539–68551.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Advances in Neural Information Processing Systems 36 (2023), 68539–68551

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:30.992778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:37:30.416102Z digest=sha256:3aa94a6845dd1f40ab879277e9d8d9ab1af89e9c091667a4a4ccc571abc312e1

Observation 640b320e-fa86-4451-90c8-29f6488264bd · outbound

This paper cites ALJP: An Arabic Legal Judgment Prediction in Personal Status Cases Using Machine Learning Models.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning ALJP: An Arabic Legal Judgment Prediction in Personal Status Cases Using Machine Learning Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:37:30.687681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:37:30.462054Z digest=sha256:410ae1430325cc0862b97640ba941b7c3b9c20c00fa4d0e65a5d6ac60d818aae

Observation 366625d7-fecf-4096-be1e-d31814a019e7 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:30.520795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:30.520795Z digest=sha256:e8697a7ba7152f540eacc50c1eeea48961a7c641e1568c1a5b60000427eaad79

Observation 02b4070b-e6d0-4dc9-9a42-fcc2995941ff · outbound

This paper cites In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers).

Reinforcement learning fine-tuning of language model for instruction following and math reasoning In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers)

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:37:31.172059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T04:37:29.669550Z digest=sha256:2dfa7f23025f6480caab3856912616671272ef4b446c4e48d97969d80b32b19e

Observation 5a2ca701-2b9b-4b81-b3e4-abec734aa4bc · outbound

This paper cites DeBERTa: Decoding-enhanced BERT with Disentangled Attention.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning DeBERTa: Decoding-enhanced BERT with Disentangled Attention

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.907111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.907111Z digest=sha256:86e57b2614dccc415d06adf6abd0d68d71283f2a63ab01922481a2838201fe09

Observation 787ecde3-dc8e-4df0-b6c3-4c3626bb138b · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.608840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.608840Z digest=sha256:a9247cf4b309373dfa6bac4b812f8198862eb488542a32695369462ca94b7274

Observation 5a339a03-76b0-496e-9880-10cae60cd3d1 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:29.438985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:29.438985Z digest=sha256:f689e3f5c68b47f23b968d7130019bb73959cb233f90eb441f5278f116da5126

Observation 072ae991-3bc8-4471-9ed3-6a86b39a3cf6 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Reinforcement learning fine-tuning of language model for instruction following and math reasoning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:37:30.121543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:37:30.121543Z digest=sha256:2879c3451fa90a7166dbf5f28495a36898343b18b80bf853a18bcfdff24921ee

Pith citing papers

No inbound Pith citation observations are available.