Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 43 inbound Pith citation observations for arXiv:2401.08967.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T10:25:13.011848Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 2b60958c-1ddf-4199-a784-b73d877bbe4e · inbound
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs ReFT: Reasoning with Reinforced Fine-Tuning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1c805dae-c164-4439-9298-0f08e61d317c · inbound
Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance ReFT: Reasoning with Reinforced Fine-Tuning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3565846-792b-4bbf-8a49-6bc802995dec · inbound
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization ReFT: Reasoning with Reinforced Fine-Tuning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b43b2d88-3886-4512-aada-45ba6ff87023 · inbound
A Survey of Scaling in Large Language Model Reasoning ReFT: Reasoning with Reinforced Fine-Tuning
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c0a4906-8f08-492a-a834-58cccecf0722 · inbound
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs ReFT: Reasoning with Reinforced Fine-Tuning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f422ba67-dcdb-42d1-abdb-77fbde5d2ca9 · inbound
The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning ReFT: Reasoning with Reinforced Fine-Tuning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b77e8eb1-cffd-4a6e-8c4c-dc90b07f2190 · inbound
Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence ReFT: Reasoning with Reinforced Fine-Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d478c978-2271-4eef-acb5-47450fea440c · inbound
ManiTaskGen: A Comprehensive Task Generator for Benchmarking and Improving Vision-Language Agents on Embodied Decision-Making ReFT: Reasoning with Reinforced Fine-Tuning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f82f1d-90bb-424d-ad69-2b7662b4c4cf · inbound
Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning ReFT: Reasoning with Reinforced Fine-Tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59a5836c-04db-4516-af7a-2d9bb671b982 · inbound
Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise ReFT: Reasoning with Reinforced Fine-Tuning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2cb4f65-b5b7-41c3-ab44-f1004e6de99a · inbound
Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning ReFT: Reasoning with Reinforced Fine-Tuning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f1138a9-c49d-4495-b9d6-6c8355ea4e03 · inbound
AI Agent Behavioral Science ReFT: Reasoning with Reinforced Fine-Tuning
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc871d76-c04e-4477-881b-39ea09382062 · inbound
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models ReFT: Reasoning with Reinforced Fine-Tuning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1ee18de-6c4a-4f6f-8ec9-ad19fbc7fa63 · inbound
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning ReFT: Reasoning with Reinforced Fine-Tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28faa975-416c-4b4a-abf9-c8e60afd8a80 · inbound
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset ReFT: Reasoning with Reinforced Fine-Tuning
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 216caaee-f178-4204-b854-b162cec63e98 · inbound
Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training ReFT: Reasoning with Reinforced Fine-Tuning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f67dfa6e-fe8d-43cf-80b9-d1ddee10220c · inbound
One Token to Fool LLM-as-a-Judge ReFT: Reasoning with Reinforced Fine-Tuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60eaf0db-3db8-40e8-92d3-e5fe0f7f32dc · inbound
Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding ReFT: Reasoning with Reinforced Fine-Tuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85fb2d14-dd27-47b3-8cea-2c384eabee1f · inbound
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice ReFT: Reasoning with Reinforced Fine-Tuning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03bae2fa-c3cb-4423-8c12-78049f450a13 · inbound
AdsQA: Towards Advertisement Video Understanding ReFT: Reasoning with Reinforced Fine-Tuning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8318f854-3f94-46a9-8799-50700ba725a1 · inbound
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards ReFT: Reasoning with Reinforced Fine-Tuning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 011d2cef-7b5a-473a-aeec-aab241eaa0a9 · inbound
VIDEOP2R: Video Understanding from Perception to Reasoning ReFT: Reasoning with Reinforced Fine-Tuning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 393c61a2-8e87-45aa-b99a-acdb263710d5 · inbound
CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical Reasoning ReFT: Reasoning with Reinforced Fine-Tuning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 134200e4-b179-45bb-83fb-61d63dfb2635 · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning ReFT: Reasoning with Reinforced Fine-Tuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 60924a86-6c78-4322-9b77-3f5f1ec98ea8 · inbound
rePIRL: Learn PRM with Inverse RL for LLM Reasoning ReFT: Reasoning with Reinforced Fine-Tuning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3931b483-e967-4b6c-b3d4-ad9276220cd8 · inbound
Towards Sparse Video Understanding and Reasoning ReFT: Reasoning with Reinforced Fine-Tuning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b20dc4d0-0773-412b-891f-83cb66077905 · inbound
PARM: Pipeline-Adapted Reward Model ReFT: Reasoning with Reinforced Fine-Tuning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e6d25f51-1e15-4dbf-b6f7-2d39a11ab1ef · inbound
Navigating the Clutter: Waypoint-Based Bi-Level Planning for Multi-Robot Systems ReFT: Reasoning with Reinforced Fine-Tuning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 604eab8c-9026-409f-8c93-f7b546795d20 · inbound
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key ReFT: Reasoning with Reinforced Fine-Tuning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 39b51824-bed8-48e7-b49f-e7b227ea4e8d · inbound
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key ReFT: Reasoning with Reinforced Fine-Tuning
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 97a3e1c7-6394-4eab-bafe-90c433301b34 · inbound
Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Key ReFT: Reasoning with Reinforced Fine-Tuning
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a5398c4e-c9e3-44eb-9592-713551c87f45 · inbound
OracleTSC: Oracle-Informed Reward Hurdle and Uncertainty Regularization for Traffic Signal Control ReFT: Reasoning with Reinforced Fine-Tuning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9b569002-79a6-411e-82d6-f71c6a131152 · inbound
Reinforcing Multimodal Reasoning Against Visual Degradation ReFT: Reasoning with Reinforced Fine-Tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 12f10086-267a-46f8-b0c7-659981a5e0d8 · inbound
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media ReFT: Reasoning with Reinforced Fine-Tuning
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 225262dd-f8f8-4b10-9332-b4e643357c8a · inbound
Hide to Guide: Learning via Semantic Masking ReFT: Reasoning with Reinforced Fine-Tuning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 52332038-93eb-4e0f-ad74-121870f0626d · inbound
Sakana Fugu Technical Report ReFT: Reasoning with Reinforced Fine-Tuning
Reference 290
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fe614fb8-4ffc-4679-b762-05bf33793149 · inbound
Improving Reasoning in Vision-Language Models via Perception Verified Self-Training ReFT: Reasoning with Reinforced Fine-Tuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 195b7c74-96db-4177-97df-00a66ee514ca · inbound
Improving Reasoning in Vision-Language Models via Perception Verified Self-Training ReFT: Reasoning with Reinforced Fine-Tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a521c81a-e063-4479-9afa-ae5901e356de · inbound
SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text ReFT: Reasoning with Reinforced Fine-Tuning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44030fd9-7c7c-4891-a80e-b818b9ea07ef · inbound
LeAct: Learning to Reason from Expert Actions ReFT: Reasoning with Reinforced Fine-Tuning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfcfd0d3-334c-4196-9e74-1351f190d653 · inbound
MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation ReFT: Reasoning with Reinforced Fine-Tuning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f2b05e2-0da4-4382-99b5-a5ede3c9c0dd · inbound
Self-Improving Large Language Models via Progressive Experience Evolution ReFT: Reasoning with Reinforced Fine-Tuning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd8b029b-c614-420d-963a-c6901beb1782 · inbound
Self-Improving Large Language Models via Progressive Experience Evolution ReFT: Reasoning with Reinforced Fine-Tuning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.