Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:23:27.192432Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2608.08224.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:23:27.192432Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a055507b-9cc2-4844-a9cf-fb1774a5aa21 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Advances in Neural Information Processing Systems (NeurIPS) , year=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09c720e4-4280-4d93-a8f4-76a21bb0f01c · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Advances in Neural Information Processing Systems (NeurIPS) , year=
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5772d3d2-5b68-4017-a57a-0ea646db5f5c · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c0a9770-474b-407e-9c5c-14dab82bd794 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb2c0d45-5c10-44cf-a228-dfb3972a0ee7 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22bcd684-565c-4e41-8d61-2288361cbb08 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Advances in Neural Information Processing Systems (NeurIPS) , year=
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0b41890b-103d-4d18-bc7f-3367627a6fef · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training and Liu, Alisa and Dziri, Nouha and Lyu, Shane and others , journal=
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f722d4af-9da7-437c-8511-7ce41877fd62 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training 2025 , howpublished=
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9dbc91ec-8d30-4a85-b75b-e4f0c38f2c53 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training 2025 , howpublished=
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c499c69-1e6a-4d53-a4d7-96b1c097a1c8 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d8c7570-6afc-4257-a1ac-c80131a63690 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a5937d4-b8bf-4ee2-aa04-5d674c99f2c2 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Proceedings of the 7th BlackboxNLP Workshop (EMNLP) , year=
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fed65094-a888-4d86-afe9-656b5468259c · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training 2021 , howpublished=
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8a92781-348e-4b16-ab7a-85198e2ba2b4 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training 2023 , howpublished=
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1ff3eea1-771f-4546-982b-10489d92c7f5 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Advances in Neural Information Processing Systems (NeurIPS) , year=
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed7e3211-8d88-4d3a-a7ad-dc791b4dd4e6 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Interpretability in the Wild: a Circuit for Indirect Object Identification in
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99ffb5c0-1972-4a93-9f01-27fb72ec9e53 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 929fb0ac-c46e-42ab-b4af-5b9dbcef26c6 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training International Conference on Machine Learning (ICML) , year=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0da3dfd1-fdf3-449d-8e36-71e9348a3295 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Conference on Language Modeling (COLM) , year=
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c1ae0d81-43fa-47ad-b972-961ca6787fdc · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training International Conference on Learning Representations (ICLR) , year=
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 40aa5ad5-69f2-477b-b54b-c4e866773d13 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training International Conference on Learning Representations (ICLR) , year=
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 95492037-5286-4823-8135-0a378bec20da · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training International Conference on Learning Representations (ICLR) , year=
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 058f235b-cfc3-438f-83cf-8f2350d2fd1d · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Biochemical Journal , volume=
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e13722e8-4788-4939-99e7-6d1074321f8b · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Symposia of the Society for Experimental Biology , volume=
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 53fea1a3-e560-45cb-96f3-d204d2cfea1a · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training European Journal of Biochemistry , volume=
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 224f42e4-896b-4004-aaae-8cac7f066e53 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Advances in Neural Information Processing Systems (NeurIPS) , year=
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fe57f195-1991-42d7-8e57-dddc10f14c5b · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training International Conference on Learning Representations (ICLR) , year=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afcf2de7-4c10-4195-9f03-ad7185b12932 · outbound
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 01d8e8d1-7353-4217-bbae-6cf437d6c40a · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 220f296b-cbe5-4803-8af7-48e389d3ba3d · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 075ce667-b95c-41ae-bc44-1f86246c8002 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Evaluating Large Language Models Trained on Code
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37a31f4b-bfc3-48aa-9dc2-10cd029b6215 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Measuring Mathematical Problem Solving with the
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a94e6974-09dc-4b80-9d13-f51a57224350 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Advances in Neural Information Processing Systems (NeurIPS) , year=
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9fdf7d35-340a-4d6a-8cf9-a938314e233e · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Program Synthesis with Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de6ecf71-5e42-44d0-a878-ea08ad374705 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Is Your Code Generated by
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5ebbaaee-44c4-4e00-85b0-e1a2c16fc801 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Proximal Policy Optimization Algorithms
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4815e24e-f3ed-4b71-85e6-06cec19a1fdc · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training and Ermon, Stefano and Rudra, Atri and R
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e1b7bc2a-7a62-4734-a5f1-e4910240fe1f · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 596ff04d-1672-46f6-aabb-f8cd0c9bef7b · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eaa5a27-8834-4c94-9767-111c41ff01fa · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training Qwen2.5 Technical Report
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 283f2359-6c66-4766-8286-f628810b91a9 · outbound
Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training The Llama 3 Herd of Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.