Pith. sign in

Paper Citation Record · LEDGER

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 17 inbound Pith citation observations for arXiv:2509.07980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07980 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:28:52.581372Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:36:25.998787Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:56:30.028956Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1cbdb3f4-4df4-4b74-b593-063bbb42ed12 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.417972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.417972Z digest=sha256:5c0d97352020d025bec4e88c904605c636d1df3107a5ce6e27c4dd0c69f23382

Observation ed11c2c4-520a-4729-99d2-62093fe0dd0c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.842982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.842982Z digest=sha256:24c839a018a3885db7096ae7786e19261ef197f43d7ae2088784642559cdfa07

Observation b198ef7d-3422-4d4b-8989-f7f26f6cbc2a · outbound

This paper cites Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context Learning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:28:53.391615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:28:49.059722Z digest=sha256:dfd1aa4504565b55bf59cb2d65621d57d12e7d1c3440399933391b53617be97e

Observation d60036a1-5e84-4f5c-a3d6-b9a9c62b5162 · outbound

This paper cites Efficient Test-Time Scaling via Self-Calibration.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Efficient Test-Time Scaling via Self-Calibration

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.143122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.143122Z digest=sha256:691c0a600629310623087101445c22bfdd2a0f71dd35f63ff3caaf54779a758f

Observation af4ec024-5037-42b1-81e7-713b7e716b3d · outbound

This paper cites Self-Rewarding Vision-Language Model via Reasoning Decomposition.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.329292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.329292Z digest=sha256:d5d0004d8c6adb79b1887705563b217bbf4a1a7ae0d2814cb5d62f0dba645788

Observation 9547c6ec-a53a-4d31-bff0-449c73ccfa24 · outbound

This paper cites Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.390681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.390681Z digest=sha256:d8b2e5c834f742c6206954a965c041ec48a872c9b9523d6dffa2f2a372d63547

Observation 80287147-8dee-4a21-9caa-501098559db8 · outbound

This paper cites Matthew Macfarlane, Minseon Kim, Nebojsa Jojic, Weijia Xu, Lucas Caccia, Xingdi Yuan, Wanru Zhao, Zhengyan Shi, and Alessandro Sordoni.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Matthew Macfarlane, Minseon Kim, Nebojsa Jojic, Weijia Xu, Lucas Caccia, Xingdi Yuan, Wanru Zhao, Zhengyan Shi, and Alessandro Sordoni

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:28:53.693150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:28:49.445211Z digest=sha256:142dcbc5ff3a8dd8ba05c2e66bca1ef43d4db315e5ca0465ccb607751326c2ce

Observation 6fc2d1c6-c9bc-49c0-ace7-5ab8591c63f1 · outbound

This paper cites Learning Adaptive Parallel Reasoning with Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning Adaptive Parallel Reasoning with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.523071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.523071Z digest=sha256:ef9996d2ec0675af7c072e143ea0b1802c859de60ab1cfb4d45176522c6028f8

Observation 713939a4-7a0f-458e-ad97-07f78449db3c · outbound

This paper cites Hogwild! inference: Parallel llm generation via concurrent attention.arXiv preprint arXiv:2504.06261,.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Hogwild! inference: Parallel llm generation via concurrent attention.arXiv preprint arXiv:2504.06261,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.606573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.606573Z digest=sha256:0fea3b2becc53ef93a2ba216f7d9d67688fdbf9855b5bffd577080333cdfedaa

Observation 405ed9da-be08-441c-a7c5-f99b938dd518 · outbound

This paper cites Adversarial Reasoning at Jailbreaking Time.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Adversarial Reasoning at Jailbreaking Time

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.697568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.697568Z digest=sha256:f0b66300f0442d1a9132d7dec929512357d98d3616161543a03f648847a097cd

Observation e08612b1-25fd-4410-b5dd-12cf2946f872 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.771954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.771954Z digest=sha256:73722e0f45d629f69e992cfa0f215c37216e0593ec64118d3088228191bacc44

Observation 92990e46-971a-4b49-802d-bc35ea597896 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.856838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.856838Z digest=sha256:dd05fe2c41645cb62fe712665b97590d82c57262c05c2ae92e3569ec8f2ef652

Observation 3a8e54aa-775a-42e5-89c2-6ac17b513180 · outbound

This paper cites MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.949465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.949465Z digest=sha256:b2b47c4b0dd996b6fb6e6346603d6b500f2efde8409a4616068667eb82e431e1

Observation bf208120-e27a-4f5a-a3f3-df5416e19825 · outbound

This paper cites On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.022641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.022641Z digest=sha256:3cb80828cafcb3ae253dc5b53a3646022e6b064ce5529c58308e74df0ec97925

Observation bf8f7d3c-d7c9-40f7-b544-a1fa7f6a0420 · outbound

This paper cites To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.169747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.169747Z digest=sha256:a0a718b3b31c55624a1a185a6904532ab19c6956e5c9cd70eddfc3022511416b

Observation 9691f506-e6da-4df4-b0aa-c82e74f335aa · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.583503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.583503Z digest=sha256:1eae1a49ced95b4c19e286bb1fee8e062721c0865e5def488df2cc2ddd9b7b46

Observation 7449f66d-9235-43f2-b462-499ca818af64 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:51.031284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:51.031284Z digest=sha256:1ea1798f9abd2e237b40d0c3ba06def6f1643614d0dd056c2e8db794f5f0e608

Observation 2128fd9f-9219-4868-92bb-21e14ce554b4 · outbound

This paper cites Learning to Reason via Mixture-of-Thought for Logical Reasoning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning to Reason via Mixture-of-Thought for Logical Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:52.055915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:52.055915Z digest=sha256:e633b6fa791852545003b0bd11ec7a70f3726976dd68276e080fe1b33eb02168

Observation 42cfe4e3-fc94-478f-9ca9-40b95fdf9932 · outbound

This paper cites Dissecting logical reasoning in llms: A fine-grained evaluation and supervision study.arXiv preprint arXiv:2506.04810,.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Dissecting logical reasoning in llms: A fine-grained evaluation and supervision study.arXiv preprint arXiv:2506.04810,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:52.354156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:52.354156Z digest=sha256:8f64c220b799b515cef3fc8e2b4f86990e3260fc060ae4d5ca23e57bb459dd7a

Observation 2220d7ac-1e80-4b0f-bb7a-a0aa533bc019 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning TTRL: Test-Time Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:52.581372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:52.581372Z digest=sha256:d48254f78dc58ee97738b32f4411a5f0a5c2f13837ea13d02e55118a409b0a4c

Observation fdcdd31d-64c2-4f09-9a52-63b8c5b54c8a · outbound

This paper cites Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models

Reference 1989

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.635933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.635933Z digest=sha256:0a2e14a015e5b844187a59d0533a3d66550828a5b8802715648c75631db59cee

Observation c59ba061-c2cb-49e9-8181-ebcf4f50a036 · outbound

This paper cites Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.250131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.250131Z digest=sha256:0ad03d0f595b0b5e69ecbbf848324869ffc447aa24c791e687d247b82e4cbad2

Observation d32bd04e-8be1-4b51-beb1-40f512cf0fcb · outbound

This paper cites Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.990999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.990999Z digest=sha256:3fafda00b34995310c36ec00dfeff9e024c990ae18f30032fe32445ef604d0fb

Observation 545b25e9-4311-4970-99e9-24fbce537175 · outbound

This paper cites Qwen3 Technical Report.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Qwen3 Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.298882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.298882Z digest=sha256:4e106f77492d564c95584f71996403660d7b55a246cc69b7f9a520b8719bd8aa

Observation 40ebbcef-d119-4401-8470-c8073730f54a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.365945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.365945Z digest=sha256:0c08e2823161c4873334555b0aa50ca2750067f631226f3d401f7d975e880eda

Observation bf4b7a5f-abee-4dad-b9af-69f47d76e78e · outbound

This paper cites ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMs.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.497263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.497263Z digest=sha256:f8b9d5758f2acc16f957ca6e646e010197f8156b581edf952d9cfc9886b7d997

Observation fe6b19d7-fe42-4231-b346-f3a6e7612a58 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.745667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.745667Z digest=sha256:6a547a03b2d9bbd24e2b070576446ccbc48690bbe40b48d326fb47be3c3b9ee8

Pith citing papers

Observation 4cb5c848-dc4a-4312-8813-dc690d41f70b · inbound

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics cites this paper.

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:25.998787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:36:25.998787Z digest=sha256:f71e0c7b1f507f14617f91ef2599144b48db2cf7b894f7b9eaffd8be3d514494

Observation 35545380-d09f-42e0-aa19-7af9659e144f · inbound

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning cites this paper.

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:13:47.999081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T01:11:26.411893Z digest=sha256:08a17ad618275ee890ebbef0370b45e070496807c0ba9ab7b878e8c68e1f945e

Observation 101c8fcc-0766-4ad4-937a-0f5757a1c1a7 · inbound

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents cites this paper.

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:58:31.796120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T20:56:58.770875Z digest=sha256:e1a9ae8d506d815483329fdef7d790b3103073bec837d94b619acf5c78e2ddf6

Observation 1e4f5c74-0e8f-4a1c-bbc1-d90185439540 · inbound

On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency cites this paper.

On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:40:51.382604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T10:38:09.875786Z digest=sha256:4ded9a80a831bcab703b5f5bd7d8ad6c0dde58403db5699f45d69d260cac9056

Observation 0050ce8a-a01e-4356-bc8e-2e84d34e485d · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:b2c3c7ed6f6b5af950677571f36f749bb463941afb3f8e55ffb55d6dd8581b87

Observation 1a452e8f-2252-4f76-a376-3f16b8a823dc · inbound

Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models cites this paper.

Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:20:52.481619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:36:01.200412Z digest=sha256:1c9880d6d067c44001c055eb07b31c4a57f2b3b7bc9a30ac039cfc88258f7e19

Observation d33036cc-7cf4-4082-92ae-3eff1461b24f · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:21.243161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T10:24:16.283375Z digest=sha256:870080a2c1f753baa62006920139808353ed2bf4ecde8dc5052cfe3d28282b10

Observation 87bd6cf4-27ac-4581-98b7-e984461203ab · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:50.152538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T00:47:51.440441Z digest=sha256:9e977d463f628451f462d261b299c1c5b1d1526a41e5247344a1d7b29720775f

Observation 27be7139-36e9-413a-bbc7-69b562b1b901 · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:28.484141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T04:05:35.063305Z digest=sha256:5d9bada34b8c9bcdedc08e30a6f1f2568e4f362dd1a21aec8a3551c85e637b7b

Observation 4bf022af-6523-4e8d-b7b2-2929d433f2d5 · inbound

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data cites this paper.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:02.009207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:0b0f73c079ee2633ce32e6cf4351373ba5d0931f337a0e5a491e6c8fbfb20104

Observation cf37284b-4403-415a-868c-d46f9d7fc1e8 · inbound

The Scaling Properties of Implicit Deductive Reasoning in Transformers cites this paper.

The Scaling Properties of Implicit Deductive Reasoning in Transformers Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:56:06.362781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T16:56:53.056292Z digest=sha256:cbbf9f87327e8600a99c3b94b800ae5c5a02fb0cf27ea6c0f091fae60b906bca

Observation fd7d90a0-e220-4ac0-9b16-9863a5e3d58e · inbound

The Scaling Properties of Implicit Deductive Reasoning in Transformers cites this paper.

The Scaling Properties of Implicit Deductive Reasoning in Transformers Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T14:54:16.625002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:54:16.625002Z digest=sha256:e546485d9faa40c7dc378e6130cea2b1205799db3e7c61c8af794c748e1bd9f6

Observation 6dc354b1-2c7a-42b3-823a-96c3d3d4d9f8 · inbound

Regulating Branch Parallelism in LLM Serving cites this paper.

Regulating Branch Parallelism in LLM Serving Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:58.350558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:00:53.308946Z digest=sha256:f8190d2e89343e604a44f3c02c5f21f87cdd5cdb46a16bea11c94757ebfd398d

Observation d9179099-d64b-4beb-9db5-5f92b6033424 · inbound

Reinforcing Multimodal Reasoning Against Visual Degradation cites this paper.

Reinforcing Multimodal Reasoning Against Visual Degradation Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:25.161203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:37:21.451146Z digest=sha256:c1c03b62e0d4a748ec9a5468a9e365ffeb26da276e7166f0ab5c8f59a8399e9f

Observation cca0ac9b-fc25-41ba-a2b9-ef19002ac348 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.337093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:9be17f3dc0e7a6057cd009523d498a111643babfa7eda2a8b4bd0ef0ad1931d1

Observation 0aeec0c5-f4f8-4218-be30-738bdc2b06b9 · inbound

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling cites this paper.

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:30.030900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T10:25:10.559953Z digest=sha256:f852bb9fd65fa6f1528c3737a7e3c674d1fc0fe15d15ed3c60f9353d48afdaee

Observation a49cfe8a-2894-4485-ab0b-ea0df3835026 · inbound

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning cites this paper.

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-02T06:37:39.525266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:37:39.525266Z digest=sha256:7ed07616b729fa64a0ab9cfffa0b68c76e08f1723aea2f10151a331ad3b6c68d