Pith. sign in

Paper Citation Record · LEDGER

An Empirical Study on Eliciting and Improving R1-like Reasoning Models

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2503.04548.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.04548 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:02:42.759310Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:36.567324Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 58656352-dd46-4280-83b0-75243c5d2710 · inbound

LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation cites this paper.

LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T13:02:23.397708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:02:23.397708Z digest=sha256:c1412dc9d61d5d6924671638eaf883767b8c45bfb1436ff76e422b174cda7e2f

Observation f33de9cf-eec2-4453-83db-1f09bebf594e · inbound

DAPO: An Open-Source LLM Reinforcement Learning System at Scale cites this paper.

DAPO: An Open-Source LLM Reinforcement Learning System at Scale An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:35:13.444477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T23:33:10.824995Z digest=sha256:8009cbcc840306ab6101e390480f923abf365e2c37600310573fc73af482e93f

Observation 06469bab-0143-4fcb-82dd-0f9ab749777c · inbound

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles cites this paper.

OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:59:03.369257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T06:59:03.112252Z digest=sha256:7c060a1ef026a3786f4b18b7b284a63e15b94eeaedb155043539ce0829ba3f52

Observation 76b6808a-bca6-4a3f-96b9-3f5d8dfea00a · inbound

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs cites this paper.

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:42:39.082992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T18:42:39.023650Z digest=sha256:3c756e0b91a07adaf9eb19828b1e3aeedb24630ee18620c6444c8a75098916f1

Observation 4381c426-e715-40cb-b99e-6afc554e7039 · inbound

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning cites this paper.

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:46:56.774430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T18:46:11.571831Z digest=sha256:5f5f51f3350bc376a89354bad88124b157f7872d19b23285b9eff7564de9fcdf

Observation b59bb1c4-1110-49ca-b5cf-2e347eb94675 · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:42.759310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:42.759310Z digest=sha256:2d97e4ff2e699db22eef551476f6eff72f4aff46084b88d043e4f8051d462459

Observation cba7db72-ecca-47c0-ab50-daa341a04b18 · inbound

SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM cites this paper.

SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:55:52.288360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:55:52.288360Z digest=sha256:83d8b288ba74d45353df88ba2513d76c1571d869bc6a2960bcc8ab1e42996e4e

Observation 29297c44-ea38-4cf0-8954-a3cb1074da32 · inbound

WebThinker: Empowering Large Reasoning Models with Deep Research Capability cites this paper.

WebThinker: Empowering Large Reasoning Models with Deep Research Capability An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:14:25.443648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T19:14:25.283645Z digest=sha256:358003f3a6296118121049e6f95b7f9bccd91aabaa4ecb32e6bbbfe79bd7ffc8

Observation fbfe2200-79d0-4f9b-9cbc-d99b11d55d45 · inbound

A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law cites this paper.

A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:48:17.354301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:48:17.354301Z digest=sha256:275119f6168e70745210945e70490040a820c544c0698867830c6cd46a93b60b

Observation 1335672c-77d4-48c8-8684-31cb153b6295 · inbound

Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models cites this paper.

Long-Short Chain-of-Thought Mixture Supervised Fine-Tuning Eliciting Efficient Reasoning in Large Language Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:56:46.208553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:56:46.208553Z digest=sha256:abfff2ca9797c1f79357578bff8dd219807900495a8e6537a4fb380dfbadca6a

Observation d42b39d2-6305-404b-bafe-abd1e7e22b27 · inbound

Learning from Peers in Reasoning Models cites this paper.

Learning from Peers in Reasoning Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:13:07.983704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:13:07.983704Z digest=sha256:90dbcd20387bcc42343a4cdb1b1a35ca49d17702587d07852b5f60365c5a5e55

Observation 1026facc-1e92-45cb-be99-89ba8001c0c8 · inbound

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models cites this paper.

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:20.985769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:20.985769Z digest=sha256:bfdba116e65d8ed2d85220668e3f59cd8d903d2e49cc8c243a8ef0a166d5ef43

Observation 1b8875ab-b37b-43d1-add0-a48501d7948a · inbound

Prior Prompt Engineering for Reinforcement Fine-Tuning cites this paper.

Prior Prompt Engineering for Reinforcement Fine-Tuning An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 1901

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:00.380468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:00.380468Z digest=sha256:8daacfa4791e71b7966720351a46550c5ba86362a485854674ba9f34456e71b4

Observation a6f741f9-565d-4385-8f11-ea6c15b61ef1 · inbound

GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents cites this paper.

GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:16:01.048333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:16:01.048333Z digest=sha256:3f30285f7de9936ac153bcd11ec049d964e94d94c550bd847a494b4737396626

Observation c5f52244-f3d1-4478-92df-621630b3f6d6 · inbound

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning cites this paper.

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:03.057637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:03.057637Z digest=sha256:541191fa4f2f74c7412a6cea44ecf70916de07b88974a474507d5b300f8e0c76

Observation fe80a669-587b-42f0-83c4-ecb28f4aebd1 · inbound

DeepRec: Towards a Deep Dive Into the Item Space with Large Language Model Based Recommendation cites this paper.

DeepRec: Towards a Deep Dive Into the Item Space with Large Language Model Based Recommendation An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:12.795266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:12.795266Z digest=sha256:9018bf4e6050735ef5fbd58b8cb2767a0bcb4cdce10e9391fbc0bf678cb6599e

Observation 198ecfa4-6606-4742-af44-4607ce100a41 · inbound

LARES: Latent Reasoning for Sequential Recommendation cites this paper.

LARES: Latent Reasoning for Sequential Recommendation An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:59:15.857633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:59:15.857633Z digest=sha256:d6de5751c7a9ddfd7f724f7f2d0ec229be1dc923c97d542cbaa7ba2c0b4a2f4d

Observation 622a08fe-2874-4516-9be6-0009b06f3410 · inbound

Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning cites this paper.

Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:16.258476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:45:16.258476Z digest=sha256:21650524261b7ca29521996ff04f1e5b57aad0c2b9ae6f4a8495a69ef6eaaeb9

Observation 4b20eca8-405a-43c7-b82f-25ee47b0b245 · inbound

How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning cites this paper.

How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:21.206536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:21.206536Z digest=sha256:02aa4d3b1357f86cea89a752f1c9f90f7ce584d864af022c12e9feefddd4cead

Observation c1dd7f1c-5a02-496d-920e-2483e7447a1f · inbound

Towards Effective Code-Integrated Reasoning cites this paper.

Towards Effective Code-Integrated Reasoning An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:25:26.696704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:25:26.696704Z digest=sha256:983955f000acd5e35cb1c702d3b17eb916f462ac2a8bc8875d133e5fc9bc7373

Observation 91706d69-cc3c-4652-96bf-4e1f55e08da2 · inbound

ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests cites this paper.

ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:02.142063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:02.142063Z digest=sha256:3c560c94e3d57f831fba8b0b84a5ccb87716aadea72791118040af000c9f55c2

Observation a14cc001-f48a-4c0f-801b-fff9e23f21f7 · inbound

Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency cites this paper.

Wait, We Don't Need to "Wait"! Removing Thinking Tokens Improves Reasoning Efficiency An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:15.802751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:19:15.802751Z digest=sha256:e92a32e934e6e1fce7acbc710b605cfc3e07c05c77e9e6ca86ae74cc28f1f7dd

Observation f7a66c95-c231-4ac5-800a-9fb4f645469d · inbound

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning cites this paper.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:42.527494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:42.527494Z digest=sha256:3febcea3d89269eac481ae56bb75f2a5203f99678190933a6533e8b054be9a42

Observation 66a3de8d-8fc8-4511-9668-9911376d2011 · inbound

CoRT: Code-integrated Reasoning within Thinking cites this paper.

CoRT: Code-integrated Reasoning within Thinking An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:46:20.731461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:46:20.731461Z digest=sha256:19ba170c54e9f0883a6aaa787a533960c674d10b9b0154b5a31c6f7ab798c41d

Observation c297cfb4-f2b0-4ecc-91df-2dc7fee1cf47 · inbound

Act-With-Think: Chunk Auto-Regressive Modeling for Generative Recommendation cites this paper.

Act-With-Think: Chunk Auto-Regressive Modeling for Generative Recommendation An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:40:48.214783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:40:48.214783Z digest=sha256:25d2084f0434fb9f34be5fa1508e8745105d7f87284239f89c2806acda005d15

Observation 6d6b4f97-6902-456c-ba54-d30bb5150131 · inbound

Reasoning-Driven Retrosynthesis Prediction with Large Language Models via Reinforcement Learning cites this paper.

Reasoning-Driven Retrosynthesis Prediction with Large Language Models via Reinforcement Learning An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T14:55:03.899694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:55:03.899694Z digest=sha256:e1388cc9cc7d89451ada1dff0a15974a40761118a6725a31c71c7ca8cd1f9f48

Observation 9a097a0b-09ff-4ebf-8733-eafab4a5893d · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:05.350521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:05.350521Z digest=sha256:63c5bc4ea9325b05fe1591755236c771d134590f7898d91cb5e41cc245476c3f

Observation 0f1e3433-c858-49cf-852f-f44ed4fbd56d · inbound

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought cites this paper.

Why Does Reasoning Length Converge? Unveiling the Underfitting-Overfitting Trade-off in Chain-of-Thought An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:31:37.763822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:31:37.763822Z digest=sha256:9d42083a7d48d184f7cb2bca6e813da367c6f4633bde631e181c3a977826fb55

Observation b17cbcda-a229-4ac6-a328-5428da7d0df3 · inbound

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework cites this paper.

Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T05:45:02.361736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:45:02.361736Z digest=sha256:49266d6aec6a6cc09aba87931a627a43f706a2e6b97ee4cf073404ef8045bb9f

Observation 5c756cda-6e36-440b-b588-c5ad41ddeee5 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.155358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:824d4dd727c2908ba2e2c3f03681a7db094b3cad972e00af4425bdf31a7bb70a

Observation 628b4a99-10af-4df0-843a-96075de8680e · inbound

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards cites this paper.

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:26:28.245340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T14:24:48.666197Z digest=sha256:89a8df50d2a8c0e6752fcb8b9ee57de9e2bec7dcf1fdaa162fb428267151f510

Observation 7a135bed-7053-4874-aa72-a95d0b31e19b · inbound

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards cites this paper.

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:48:29.373053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:48:29.373053Z digest=sha256:4c4139d72f26932f7fa818a868c9ffaebc67400f9b6dbddd1699541429c2262f

Observation f6ece2d3-1975-4510-b898-b35c9b2453fa · inbound

EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget cites this paper.

EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:56:08.702299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T08:53:31.803096Z digest=sha256:de820902a666c4cb5e4c7c1c290749e03b6dc149d33f7f1ca43f24fdf9b01e71

Observation a9173ba0-c924-44b1-b186-7d8790d0b7f9 · inbound

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors cites this paper.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.275213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:fc77e489ba6d332517f44cd0c5ca69869dafed7ad6d28b64fed86188ddd214a0

Observation 6e29e227-7d4b-4b70-bf5f-5f7909f88ea4 · inbound

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning cites this paper.

PruneTIR: Inference-Time Tool Call Pruning for Effective yet Efficient Tool-Integrated Reasoning An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:01:23.360130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-12T04:42:49.165066Z digest=sha256:d67c30717a621dd0d52c165ede76f0c736ccd00062938d0df54012167ba8059a

Observation 6e0ee3e9-a6f6-4d9c-8ee7-5edacf66493d · inbound

TimelineReasoner: Advancing Timeline Summarization with Large Reasoning Models cites this paper.

TimelineReasoner: Advancing Timeline Summarization with Large Reasoning Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:01.053722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T21:33:48.376967Z digest=sha256:fcf3739f9984031f7e288c719fe58ff1824694fb9e86627066836892e3843e33

Observation 28301dbc-933c-438e-98be-bb5d6bead1b3 · inbound

SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs cites this paper.

SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:13:44.094363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-20T20:09:42.841695Z digest=sha256:9059a04a597789246bfb8363a06df37028a2e253f6957d3b37c6b670531092b9

Observation 9a52c77f-62d9-4e31-bdcd-bf21981882e7 · inbound

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data cites this paper.

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:43:55.066083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-29T19:34:12.081362Z digest=sha256:94b5d5094be92ffe2b0dd4977f68259fe4bfd78423ea8326f5042307a16ee157

Observation 02371df3-e049-457f-a350-f324403fbc61 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 289

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.161760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:d14b3a1bc3b58770c6e68cf325c13cb4d6ec3d5256f583527c1585a9e4e5d50b

Observation 6f19636a-2dcb-4fd8-9d09-89c43a4236a4 · inbound

GUI-AC: Enhancing Continual Learning in GUI Agents cites this paper.

GUI-AC: Enhancing Continual Learning in GUI Agents An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:27:36.568792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T13:56:09.049753Z digest=sha256:cad6c8d81ee81a62a489c93df9f9894be65740772b295d91d456c29a34709e59

Observation 9ea71d90-980c-4f64-9ae2-54906282b8b8 · inbound

GUI-AC: Enhancing Continual Learning in GUI Agents cites this paper.

GUI-AC: Enhancing Continual Learning in GUI Agents An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T14:27:05.589465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:27:05.589465Z digest=sha256:db75bb3574f514c34f6411aa3acfd3b07394612cad8233145416fea27a190bf8

Observation a1bd2f8b-efcb-4583-ba4e-60f18ced6894 · inbound

LaRec: Unleashing LLM-based Latent Reasoning for Generative Recommendation cites this paper.

LaRec: Unleashing LLM-based Latent Reasoning for Generative Recommendation An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:30:24.581995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:30:24.581995Z digest=sha256:c1f3e86a03c4d2952ec514b7802d8b68a6122d904b42ffa7cfb7a3b7888896f0