Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:37:30.520795Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2506.21560.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:37:30.520795Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 67b22cb6-de3c-4a24-b323-b1a84269d16e · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning arXiv preprint arXiv:2412.15287 (2024)
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4301b55-41a7-4964-a3b9-40574ff41443 · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning Stream of Search (SoS): Learning to Search in Language
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b9a162a-60ac-446a-8ddc-30c8f0d508da · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b454a4-95ff-474f-bee4-d78c48d97283 · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning Qwen2.5-Coder Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59233e2a-8955-4a6b-a21e-f05e28ca516c · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning GPT-4o System Card
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e941f99-4fb5-4202-af35-0900f6a5caa0 · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning Self-Refine: Iterative Refinement with Self-Feedback
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 259ad9ec-5d56-40a1-b865-db09449d6c89 · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ac134b5-0f06-4d92-bcf6-248a6fe79e99 · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2596d1b-9eaa-4455-a10c-3bd2053ee55f · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning Advances in Neural Information Processing Systems 36 (2023), 68539–68551
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 640b320e-fa86-4451-90c8-29f6488264bd · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning ALJP: An Arabic Legal Judgment Prediction in Personal Status Cases Using Machine Learning Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 366625d7-fecf-4096-be1e-d31814a019e7 · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02b4070b-e6d0-4dc9-9a42-fcc2995941ff · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers)
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5a2ca701-2b9b-4b81-b3e4-abec734aa4bc · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning DeBERTa: Decoding-enhanced BERT with Disentangled Attention
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 787ecde3-dc8e-4df0-b6c3-4c3626bb138b · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a339a03-76b0-496e-9880-10cae60cd3d1 · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072ae991-3bc8-4471-9ed3-6a86b39a3cf6 · outbound
Reinforcement learning fine-tuning of language model for instruction following and math reasoning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.