Pith. sign in

Paper Citation Record · LEDGER

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

As of 22 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 17 inbound Pith citation observations for arXiv:2509.07980.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.07980 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T21:28:52.581372Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:36:25.998787Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:56:30.028956Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1cbdb3f4-4df4-4b74-b593-063bbb42ed12 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.417972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.417972Z digest=sha256:117d0c0427b5b296c562e5dafdb4080f458314547effe537e511e20af7cc3376

Observation ed11c2c4-520a-4729-99d2-62093fe0dd0c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.842982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.842982Z digest=sha256:46eca2439e27942ab665f93a076acc2961582d5c21bbe7fa2c3d560e16fb6086

Observation b198ef7d-3422-4d4b-8989-f7f26f6cbc2a · outbound

This paper cites Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context Learning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Divide, Reweight, and Conquer: A Logit Arithmetic Approach for In-Context Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:28:53.391615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T21:28:49.059722Z digest=sha256:65eb0d1aac77bad87046270dc33e27812d59fcf48ae2ec0b5f3bdebe2ee3a0c9

Observation d60036a1-5e84-4f5c-a3d6-b9a9c62b5162 · outbound

This paper cites Efficient Test-Time Scaling via Self-Calibration.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Efficient Test-Time Scaling via Self-Calibration

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.143122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.143122Z digest=sha256:a7f00e69dabcf4cf314c4fad9d52e6e6b367644f53f38489e1e70f77237e7a7c

Observation af4ec024-5037-42b1-81e7-713b7e716b3d · outbound

This paper cites Self-Rewarding Vision-Language Model via Reasoning Decomposition.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.329292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.329292Z digest=sha256:34e0e43d77bc653a72e2fcf7773ef41cbb6e8c434a9024bf5786d69df401ee0c

Observation 9547c6ec-a53a-4d31-bff0-449c73ccfa24 · outbound

This paper cites Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Spiral: Self-play on zero-sum games incentivizes reasoning via multi-agent multi-turn reinforcement learning.arXiv preprint arXiv:2506.24119,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.390681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.390681Z digest=sha256:11f8621e70812db0c8f2b2e78d9245fb9bb05b8e7690e1f82d3cdcadd595b7e6

Observation 80287147-8dee-4a21-9caa-501098559db8 · outbound

This paper cites Matthew Macfarlane, Minseon Kim, Nebojsa Jojic, Weijia Xu, Lucas Caccia, Xingdi Yuan, Wanru Zhao, Zhengyan Shi, and Alessandro Sordoni.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Matthew Macfarlane, Minseon Kim, Nebojsa Jojic, Weijia Xu, Lucas Caccia, Xingdi Yuan, Wanru Zhao, Zhengyan Shi, and Alessandro Sordoni

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T21:28:53.693150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-04T21:28:49.445211Z digest=sha256:1e16553a24f7ffd084ce508830ee603404d7824afc0f819b23ffe275a9baa403

Observation 6fc2d1c6-c9bc-49c0-ace7-5ab8591c63f1 · outbound

This paper cites Learning Adaptive Parallel Reasoning with Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning Adaptive Parallel Reasoning with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.523071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.523071Z digest=sha256:2fb98e669db1fb9afae3120b5b90a7d4b60b4cc45ddcea467f1b7bba4b79a3bb

Observation 713939a4-7a0f-458e-ad97-07f78449db3c · outbound

This paper cites Hogwild! inference: Parallel llm generation via concurrent attention.arXiv preprint arXiv:2504.06261,.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Hogwild! inference: Parallel llm generation via concurrent attention.arXiv preprint arXiv:2504.06261,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.606573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.606573Z digest=sha256:b43abfe93c5e4aafd112745c7c94a96dbd03e37b5203a959b4658ba9b0be2b64

Observation 405ed9da-be08-441c-a7c5-f99b938dd518 · outbound

This paper cites Adversarial Reasoning at Jailbreaking Time.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Adversarial Reasoning at Jailbreaking Time

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.697568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.697568Z digest=sha256:3f7ef2f7e3ce4ddfe0087042774c0356cd2f108e5caa92de908afa8b093480d1

Observation e08612b1-25fd-4410-b5dd-12cf2946f872 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.771954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.771954Z digest=sha256:0cb12dd0d006144c165fc5ee5308368c831264b4e47fe51b4df421b45f0cb595

Observation 92990e46-971a-4b49-802d-bc35ea597896 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.856838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.856838Z digest=sha256:7968dc165a7d081d1a67a6a70269fd44f5278383bac59db42a8dfed787d37710

Observation 3a8e54aa-775a-42e5-89c2-6ac17b513180 · outbound

This paper cites MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.949465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.949465Z digest=sha256:f9349f11fbba879ec7535c4dfa4bce02a95424c0d3d9fc9f8f50d70fce9e3826

Observation bf208120-e27a-4f5a-a3f3-df5416e19825 · outbound

This paper cites On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.022641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.022641Z digest=sha256:504eea1a2edcd28f38268506116c8079c99f4d459c11279f37d51a0305aeb649

Observation bf8f7d3c-d7c9-40f7-b544-a1fa7f6a0420 · outbound

This paper cites To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.169747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.169747Z digest=sha256:6174fbc99a26886b43386e9854296412d1aa1f2b077e0207d3281f10819d29f9

Observation 9691f506-e6da-4df4-b0aa-c82e74f335aa · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.583503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.583503Z digest=sha256:a47a62510ca18d7c93b687c8784a145c7e3d20ee8443bece96a5aa73bbe935b3

Observation 7449f66d-9235-43f2-b462-499ca818af64 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:51.031284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:51.031284Z digest=sha256:d31a3147aa574914d91968e1fc66dd106cfba8bee36b80647266cffe2c042860

Observation 2128fd9f-9219-4868-92bb-21e14ce554b4 · outbound

This paper cites Learning to Reason via Mixture-of-Thought for Logical Reasoning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning to Reason via Mixture-of-Thought for Logical Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:52.055915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:52.055915Z digest=sha256:3a6313b1cb31eb2ec4b1245d45777a55f0577b1f4be169fc5f01e15c530155f6

Observation 42cfe4e3-fc94-478f-9ca9-40b95fdf9932 · outbound

This paper cites Dissecting logical reasoning in llms: A fine-grained evaluation and supervision study.arXiv preprint arXiv:2506.04810,.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Dissecting logical reasoning in llms: A fine-grained evaluation and supervision study.arXiv preprint arXiv:2506.04810,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:52.354156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:52.354156Z digest=sha256:1d64c61ed1dd56c292ccef3751634f0f80579290eb7e898215e03f835ced3261

Observation 2220d7ac-1e80-4b0f-bb7a-a0aa533bc019 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning TTRL: Test-Time Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:52.581372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:52.581372Z digest=sha256:08fc0195a4943379ce88999daabbe8d04a2f2fdb0b79e0177697add77f0ffe24

Observation fdcdd31d-64c2-4f09-9a52-63b8c5b54c8a · outbound

This paper cites Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Breach in the Shield: Unveiling the Vulnerabilities of Large Language Models

Reference 1989

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.635933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.635933Z digest=sha256:52dbd6801e069a2be322eebaded29b390667639f1f537a1b24097d5e7855771a

Observation c59ba061-c2cb-49e9-8181-ebcf4f50a036 · outbound

This paper cites Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Learning to Keep a Promise: Scaling Language Model Decoding Parallelism with Learned Asynchronous Decoding

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:49.250131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:49.250131Z digest=sha256:4bee7fd58932337d2e1627b7cf0f8784bcc272414d22ca345bf825a4014587c7

Observation d32bd04e-8be1-4b51-beb1-40f512cf0fcb · outbound

This paper cites Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.990999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.990999Z digest=sha256:53484fed5df7062840b30c7b1d16af3977297756cae6fe881d713159f9c5093a

Observation 545b25e9-4311-4970-99e9-24fbce537175 · outbound

This paper cites Qwen3 Technical Report.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning Qwen3 Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.298882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.298882Z digest=sha256:30fc8bdece3ac8f76864725303f70f795c1bb15ec7785ce43fce0dabb2eee391

Observation 40ebbcef-d119-4401-8470-c8073730f54a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:50.365945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:50.365945Z digest=sha256:5c25b349981b4bbed5b43a69f8cc9a64756645fe79da6277050d622e553d1a1b

Observation bf4b7a5f-abee-4dad-b9af-69f47d76e78e · outbound

This paper cites ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMs.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.497263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.497263Z digest=sha256:6fbed609ca2472eaf3630df834964075eeb314587ebc9a0df366bebe50dcde64

Observation fe6b19d7-fe42-4231-b346-f3a6e7612a58 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T21:28:48.745667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:28:48.745667Z digest=sha256:bd71aa2ae1a5e8391a23412032b9bfd79a23a77760fbfd26fd58e94da2100aa9

Pith citing papers

Observation 4cb5c848-dc4a-4312-8813-dc690d41f70b · inbound

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics cites this paper.

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:25.998787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:36:25.998787Z digest=sha256:f7a25fc7d1aaf45ea209927cccedb370d3674dccd5b1a9c578fa0531f016f8c1

Observation 35545380-d09f-42e0-aa19-7af9659e144f · inbound

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning cites this paper.

Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:13:47.999081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T01:11:26.411893Z digest=sha256:641f4f675bd9a46d5fac133a68fe38d1f21c63a5ab266600820abeb0a2eb9e1a

Observation 101c8fcc-0766-4ad4-937a-0f5757a1c1a7 · inbound

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents cites this paper.

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:58:31.796120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T20:56:58.770875Z digest=sha256:b1715b1dce8328b1bb81c1d8ab9f1dd430a631e5aa7ab0af24e05922f75f74f8

Observation 1e4f5c74-0e8f-4a1c-bbc1-d90185439540 · inbound

On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency cites this paper.

On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:40:51.382604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T10:38:09.875786Z digest=sha256:b59184ec2e6102c08c8981518b5f2b521bff1c9490ab9a798bd82e61a04ce145

Observation 0050ce8a-a01e-4356-bc8e-2e84d34e485d · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:ff059752a033c9b2f0f6593bd4e1594d4a6ea2d52baa46468204dc341ef62f72

Observation 1a452e8f-2252-4f76-a376-3f16b8a823dc · inbound

Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models cites this paper.

Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:20:52.481619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:36:01.200412Z digest=sha256:fde02535a890945c8f3d934ddf0e60f9249ba114393c3486b790708b875951d0

Observation d33036cc-7cf4-4082-92ae-3eff1461b24f · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:21.243161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T10:24:16.283375Z digest=sha256:b4ad3c9dad09c31e129688c014be411f5c33598349dfdbe946f34afac593fed8

Observation 87bd6cf4-27ac-4581-98b7-e984461203ab · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:50.152538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-11T00:47:51.440441Z digest=sha256:a7cdae2c69d8e09f50373307d17b3a6b16bc8ec7403fe52581bb0da19656c955

Observation 27be7139-36e9-413a-bbc7-69b562b1b901 · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:28.484141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-12T04:05:35.063305Z digest=sha256:ec8da7c02c53155852406ffe0f76b2ae7a9a61b325101f34e2831f822ceb473c

Observation 4bf022af-6523-4e8d-b7b2-2929d433f2d5 · inbound

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data cites this paper.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:02.009207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:2aabb3d895b2953b7560f92cbe98623936f7dac20a892fda2039d78d4042b17b

Observation cf37284b-4403-415a-868c-d46f9d7fc1e8 · inbound

The Scaling Properties of Implicit Deductive Reasoning in Transformers cites this paper.

The Scaling Properties of Implicit Deductive Reasoning in Transformers Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:56:06.362781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T16:56:53.056292Z digest=sha256:2ead9708f3d977d60b1c5fd5c304d6d7fa8ce1107735b6f9fef551efcf9f98c5

Observation fd7d90a0-e220-4ac0-9b16-9863a5e3d58e · inbound

The Scaling Properties of Implicit Deductive Reasoning in Transformers cites this paper.

The Scaling Properties of Implicit Deductive Reasoning in Transformers Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T14:54:16.625002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:54:16.625002Z digest=sha256:fee92bede31b6222ebc85481ca8091b1bb5cd7238357b96e551bb3f7325f35ce

Observation 6dc354b1-2c7a-42b3-823a-96c3d3d4d9f8 · inbound

Regulating Branch Parallelism in LLM Serving cites this paper.

Regulating Branch Parallelism in LLM Serving Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:58.350558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T01:00:53.308946Z digest=sha256:abbc452fa38ad895abd37c10c8e2dee522ee3607e81349531e63ea45f30057a1

Observation d9179099-d64b-4beb-9db5-5f92b6033424 · inbound

Reinforcing Multimodal Reasoning Against Visual Degradation cites this paper.

Reinforcing Multimodal Reasoning Against Visual Degradation Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:25.161203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:37:21.451146Z digest=sha256:977e9114147be165c687e83de9341c6505f5972e8e0559f17662cd77fc2eeafe

Observation cca0ac9b-fc25-41ba-a2b9-ef19002ac348 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.337093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:092a45c1a2409b68731a065c733a7b07687f670512b7985ddc2dfbc0c0e762d5

Observation 0aeec0c5-f4f8-4218-be30-738bdc2b06b9 · inbound

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling cites this paper.

Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:30.030900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T10:25:10.559953Z digest=sha256:815e499de2f02a3c2876280fce77baf5d092e2a548f12c84143744f8aec8179c

Observation a49cfe8a-2894-4485-ab0b-ea0df3835026 · inbound

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning cites this paper.

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-02T06:37:39.525266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:37:39.525266Z digest=sha256:d6092b8c5ace76a414c285ec72a8c9fb1f9d09866839a11b1b078bae91c4576e