Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T21:28:52.581372Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 17 inbound Pith citation observations for arXiv:2509.07980.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T21:28:52.581372Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:36:25.998787Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T02:56:30.028956Z
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1cbdb3f4-4df4-4b74-b593-063bbb42ed12 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed11c2c4-520a-4729-99d2-62093fe0dd0c · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b198ef7d-3422-4d4b-8989-f7f26f6cbc2a · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d60036a1-5e84-4f5c-a3d6-b9a9c62b5162 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Efficient Test-Time Scaling via Self-Calibration
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af4ec024-5037-42b1-81e7-713b7e716b3d · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9547c6ec-a53a-4d31-bff0-449c73ccfa24 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80287147-8dee-4a21-9caa-501098559db8 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Matthew Macfarlane, Minseon Kim, Nebojsa Jojic, Weijia Xu, Lucas Caccia, Xingdi Yuan, Wanru Zhao, Zhengyan Shi, and Alessandro Sordoni
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6fc2d1c6-c9bc-49c0-ace7-5ab8591c63f1 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning Adaptive Parallel Reasoning with Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 713939a4-7a0f-458e-ad97-07f78449db3c · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Hogwild! inference: Parallel llm generation via concurrent attention.arXiv preprint arXiv:2504.06261,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405ed9da-be08-441c-a7c5-f99b938dd518 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Adversarial Reasoning at Jailbreaking Time
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e08612b1-25fd-4410-b5dd-12cf2946f872 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92990e46-971a-4b49-802d-bc35ea597896 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a8e54aa-775a-42e5-89c2-6ac17b513180 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf208120-e27a-4f5a-a3f3-df5416e19825 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf8f7d3c-d7c9-40f7-b544-a1fa7f6a0420 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9691f506-e6da-4df4-b0aa-c82e74f335aa · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7449f66d-9235-43f2-b462-499ca818af64 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2128fd9f-9219-4868-92bb-21e14ce554b4 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning to Reason via Mixture-of-Thought for Logical Reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42cfe4e3-fc94-478f-9ca9-40b95fdf9932 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Dissecting logical reasoning in llms: A fine-grained evaluation and supervision study.arXiv preprint arXiv:2506.04810,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2220d7ac-1e80-4b0f-bb7a-a0aa533bc019 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning TTRL: Test-Time Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdcdd31d-64c2-4f09-9a52-63b8c5b54c8a · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models
Reference 1989
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c59ba061-c2cb-49e9-8181-ebcf4f50a036 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d32bd04e-8be1-4b51-beb1-40f512cf0fcb · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 545b25e9-4311-4970-99e9-24fbce537175 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Qwen3 Technical Report
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40ebbcef-d119-4401-8470-c8073730f54a · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf4b7a5f-abee-4dad-b9af-69f47d76e78e · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMs
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe6b19d7-fe42-4231-b346-f3a6e7612a58 · outbound
Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cb5c848-dc4a-4312-8813-dc690d41f70b · inbound
StatEval: A Comprehensive Benchmark for Large Language Models in Statistics Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35545380-d09f-42e0-aa19-7af9659e144f · inbound
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 101c8fcc-0766-4ad4-937a-0f5757a1c1a7 · inbound
LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e4f5c74-0e8f-4a1c-bbc1-d90185439540 · inbound
On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0050ce8a-a01e-4356-bc8e-2e84d34e485d · inbound
Efficient Reasoning on the Edge Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a452e8f-2252-4f76-a376-3f16b8a823dc · inbound
Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d33036cc-7cf4-4082-92ae-3eff1461b24f · inbound
LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87bd6cf4-27ac-4581-98b7-e984461203ab · inbound
LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27be7139-36e9-413a-bbc7-69b562b1b901 · inbound
LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bf022af-6523-4e8d-b7b2-2929d433f2d5 · inbound
Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf37284b-4403-415a-868c-d46f9d7fc1e8 · inbound
The Scaling Properties of Implicit Deductive Reasoning in Transformers Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd7d90a0-e220-4ac0-9b16-9863a5e3d58e · inbound
The Scaling Properties of Implicit Deductive Reasoning in Transformers Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dc354b1-2c7a-42b3-823a-96c3d3d4d9f8 · inbound
Regulating Branch Parallelism in LLM Serving Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9179099-d64b-4beb-9db5-5f92b6033424 · inbound
Reinforcing Multimodal Reasoning Against Visual Degradation Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cca0ac9b-fc25-41ba-a2b9-ef19002ac348 · inbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0aeec0c5-f4f8-4218-be30-738bdc2b06b9 · inbound
Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a49cfe8a-2894-4485-ab0b-ea0df3835026 · inbound
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.