Pith. sign in

Paper Citation Record · LEDGER

A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2504.11343.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.11343 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:59:21.875260Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T04:05:55.456269Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f3cfd3a9-54be-4a41-b964-576ef5d6a952 · inbound

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning cites this paper.

Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-22T18:46:57.054014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T18:46:11.571831Z digest=sha256:1355fc63f7eb0acdec5144e5032606e50adbbab94af9cf253c4216afda4a905f

Observation a861f957-35dc-4297-b7ef-48b40dc9a42e · inbound

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL cites this paper.

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:59:21.875260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:59:21.875260Z digest=sha256:3369cc1bfbc98e94ad485aa37d3179402b98dd4350c7cfb5e82dbc65167705e4

Observation 1044172e-f93b-435e-aac4-4ae269e5bcd1 · inbound

Scalable Chain of Thoughts via Elastic Reasoning cites this paper.

Scalable Chain of Thoughts via Elastic Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:13:33.633613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:13:33.633613Z digest=sha256:20b2d8f53616be441900103a237b0d20396e21bc9bcc79607d7ea77ad9fb8c67

Observation 7825dc23-4a1e-4b4f-85c6-2153f653c563 · inbound

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization cites this paper.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.040337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.040337Z digest=sha256:a697cec7d64490e18e506e7bf8d8ec40296269ebdfdd6b2d9dca2a73ca6769b4

Observation ead49eaa-8eb3-46c2-b3bd-ebb996ea348c · inbound

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs cites this paper.

Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:29:01.530228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:29:01.530228Z digest=sha256:35a3299c01200f52ca9b919238d608dd114be3ec823781061e455cc53eed9f36

Observation 007b1964-85c9-4886-b1c9-f9788c4ad213 · inbound

The Hallucination Tax of Reinforcement Finetuning cites this paper.

The Hallucination Tax of Reinforcement Finetuning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:44:35.621729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:44:35.621729Z digest=sha256:57ffa0efe618d649b4034ff4d974c05079e6733b92845090a511ad7489b373bb

Observation 55773bad-37ca-41b8-a1eb-2acc1865b435 · inbound

Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning cites this paper.

Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-22T13:34:53.196041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T13:34:27.152447Z digest=sha256:936e580997c72f60d574cd926934e5e7001ec6097523ff0e75eea30ac213096d

Observation bc56ceb1-ca1e-45ab-a9e7-534e7fc3ea92 · inbound

Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning cites this paper.

Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:35.860086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:35.860086Z digest=sha256:28fb92fb648d4caa4a1ad4aba6ef7331e5a6bd05577cd91445d80844db90a17c

Observation b7f103ad-20e3-4ede-b2fb-bba1a62b64ea · inbound

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts cites this paper.

Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:56.370913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:56.370913Z digest=sha256:eea82a03360aba92d590b303e224770ea7c290faa0ec02481cbf0bccf1f44553

Observation 962e3e5d-9e77-4b1e-aa22-9c407beb8a47 · inbound

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning cites this paper.

SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:24.239657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:24.239657Z digest=sha256:ed7ef5152dc8f1074d163554012bd47d32b8cad857e3286090c3390123070263

Observation 93597800-d7a9-4053-93fb-d358917a779b · inbound

Customizing Speech Recognition Model with Large Language Model Feedback cites this paper.

Customizing Speech Recognition Model with Large Language Model Feedback A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:21.148418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:23:21.148418Z digest=sha256:c7f53639a1bcf835906f85dc274c8a9d65af374a637424fd97120bc92acc5795

Observation 130bbb4e-c597-480b-8b5d-2e281be011ba · inbound

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation cites this paper.

Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:07:05.118325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:07:05.118325Z digest=sha256:a7ab2d936ed5fa5b08cf161b392b6671289b150f60785d446eabd3abc6efd6ff

Observation 02e451e0-3775-48b3-b6f6-f7ce19f80b58 · inbound

Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents cites this paper.

Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:05:16.165898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:05:16.165898Z digest=sha256:a202242cdb3df3010ed57bc5789dc36f63234b38821f96573643683893c8cd17

Observation dfe89364-6f1e-4824-b5d1-26b29b027228 · inbound

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model cites this paper.

AdapThink: Adaptive Thinking Preferences for Reasoning Language Model A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:40.648243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:40.648243Z digest=sha256:94b08bca4147107cbd33a91aec9c2a38ac203e9b140d06a9cb0020a7c316fd54

Observation 11c90913-72bb-4836-a257-2e2879251b27 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.253465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.253465Z digest=sha256:4badadf24a8e7158ed64c77712a1595352a2a218e11925962e29705b2d34f092

Observation 448730c8-885a-4819-9491-e48d9a58ada9 · inbound

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting cites this paper.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.525759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.525759Z digest=sha256:6ca7e1030a9a15b946d531ca95f99122978f199c9cd3042448c114895f801a02

Observation 8784bc97-94c8-4a9b-8e9f-ca944336d9f3 · inbound

AgentCDM: Enhancing Multi-Agent Collaborative Decision-Making via ACH-Inspired Structured Reasoning cites this paper.

AgentCDM: Enhancing Multi-Agent Collaborative Decision-Making via ACH-Inspired Structured Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:32:59.005705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:32:59.005705Z digest=sha256:71fa49cf9c2f9e6f1b2af914ffb146354e08f4d78da8dae832ffa2509023aada

Observation 30cb60d8-9c69-45c1-8da3-8c4b6d8d5319 · inbound

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance cites this paper.

G$^2$RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:19:46.684636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:19:46.684636Z digest=sha256:9f2d3222ac982a57768f500dbe85da96f4ef939dfe130c83380907c167ad60ac

Observation bb2d36c3-c7c9-4a0d-9b02-8641cb2cc399 · inbound

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning cites this paper.

Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:41:50.456853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T20:40:44.496392Z digest=sha256:244f7cf741e211fcc23f3d2967b4450ceeae7cc0b52ad251a6a5f56d5bc6d1cd

Observation 0c7116e4-8cdc-42c2-9271-48c094d9c972 · inbound

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training cites this paper.

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.244574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T22:38:57.833414Z digest=sha256:c4d9cfda8eb2493c3353a3b8312073b0920f1dce6809a60446c1de585ec36edf

Observation 3a903c09-9f27-48b5-996a-8758fa183abe · inbound

DiffusionNFT: Online Diffusion Reinforcement with Forward Process cites this paper.

DiffusionNFT: Online Diffusion Reinforcement with Forward Process A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:54:30.986262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T16:54:30.953199Z digest=sha256:383ab9fcf4af779590be73d2acbc2a2580082056260e6fa57c9e2b7317eaf8c3

Observation 9bb53df9-d8d6-4975-bdc3-ab161ec4c708 · inbound

Simple Policy Gradients for Reasoning with Diffusion Language Models cites this paper.

Simple Policy Gradients for Reasoning with Diffusion Language Models A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:53.871009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:53.871009Z digest=sha256:1ed9b88d964ade0c1a47a0fbc579dd0e0eedd585f0649293ab5ac76dcd416e4d

Observation 3d40d08e-cbcd-4cba-a805-8b25870ed109 · inbound

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting cites this paper.

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T08:49:34.587109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:49:34.587109Z digest=sha256:2155a7b9bdf87de9309ba50385b8cc308de16e9036615b4e073ce319f295f395

Observation e2e87cb3-0435-455e-8590-bda005b55f76 · inbound

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework cites this paper.

TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T23:33:50.580515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:33:50.580515Z digest=sha256:d9d2c4d4476b9a6f3dcb4d24a63378156d1420da0badbc97175942bdbeefa574

Observation acc3680d-d292-4990-a7da-29ed62b2cacd · inbound

Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control cites this paper.

Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:51:25.636355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T13:34:06.461850Z digest=sha256:d846d06e4a54a4254c74ca37b4e31b24ac4a1a0dd50f686524d9791e82060570

Observation 73b77fc9-59fc-469e-91af-c9034311ff2f · inbound

Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control cites this paper.

Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:32.380758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T01:56:51.940356Z digest=sha256:734aaf104facab212cdbc6d24e7bc1375a126b429fad46851030115c3290202b

Observation 1755dd5f-1fb3-48ea-8ae5-caa2491a1599 · inbound

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning cites this paper.

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:11:28.859103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-07T07:00:32.206081Z digest=sha256:74353fc3c5966f89873febcf77d4e1ec1dba1ee87506f641a995289f4ab72c12

Observation 630449dd-b80d-4aa3-a288-bfdc3eeaca58 · inbound

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning cites this paper.

Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:49:14.948536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T23:47:53.282259Z digest=sha256:3833fd036a8cbd167f4807950b84e73e2bdbf6823e0896ffc1ed5c3310970e5a

Observation 4e3f625f-f911-47dc-913c-9b7283007d79 · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:09.087260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:36e3208618f46fd2319b9ef4dd9bdcd0cf5313cdad8ae4b4dc06a18e8243f599

Observation 27f16019-f5ab-482c-a3c3-229879c8f308 · inbound

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR cites this paper.

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:10.409003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-09T15:39:54.001327Z digest=sha256:2cbb510faef144fe78162bd753e370929c7d669dea51e74429ab00d1f27f01a1

Observation 8904c37e-bae8-43f4-b1f4-e6b1b9576a75 · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.827682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:16d5c40cc327219f6265628d6317640924abab51f76609ef8a050d3ce5abe77d

Observation ac7f4f7a-08a9-4ae7-9c10-ac45239fb51b · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:55.092691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:7deff8517dfa12319167abcb9c63753be7c0c4dcebc25ae5dc19aa1a3d6d23ab

Observation cffefd4d-5467-42c3-b666-e5c677216762 · inbound

Selective Off-Policy Reference Tuning with Plan Guidance cites this paper.

Selective Off-Policy Reference Tuning with Plan Guidance A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:05.243802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T01:28:18.615371Z digest=sha256:64154e4066d0a55f92d24eed58efc8a201a833a82790b2d56d209df1f2f1b7c9

Observation 04ef7fec-55a2-469b-b11e-13f85b769347 · inbound

Selective Off-Policy Reference Tuning with Plan Guidance cites this paper.

Selective Off-Policy Reference Tuning with Plan Guidance A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:22:59.403633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T21:20:24.066520Z digest=sha256:108e55cd08bfa7132ef1eeee024055da84a739b1d3dc64e5f5c91ac6f4a5624d

Observation 6d45c1f7-88b7-4b72-9a5e-75ef342eb4fc · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:07:28.062118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T07:03:00.503644Z digest=sha256:ae4f704d9b40a03fb92598324bd9a4ec7f0ad97d51e6eaf77e93320a5e13856e

Observation f3cd6cc3-8198-42f8-b0c1-5d701390201b · inbound

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models cites this paper.

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:19:28.199683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T21:06:01.667173Z digest=sha256:b080cfe260bccd5dd002c966c996421f24adf9f8ecde942ee71be43b92f1a709

Observation fd58e815-33b8-4c2b-a799-7d7328ae7e9f · inbound

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards cites this paper.

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:59:41.063061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T05:55:45.654673Z digest=sha256:ba03166ffcd76db4ad01d6f6a4a43e2442db6829129ead344dc573b38b8cd949

Observation 0aee1525-3fe3-446e-a08f-0a4820de4161 · inbound

ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate cites this paper.

ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:00.804526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T22:54:53.415067Z digest=sha256:2b362a676e9f25314c3808a7c7129a81d2e7c31f2de112660d14e383f21f359f

Observation c20fbf8a-6285-40a2-b777-36c0e8059a36 · inbound

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection cites this paper.

Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:26:22.844157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T14:20:06.381334Z digest=sha256:f650d0ac21f0d296b2d83a34d33e9fc1859ce0a07580ccd3afc9cdb9b130760a

Observation 247f0921-085c-40d3-bb1a-5832dc6274cc · inbound

Rollout-Level Advantage-Prioritized Experience Replay for GRPO cites this paper.

Rollout-Level Advantage-Prioritized Experience Replay for GRPO A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:06:41.052519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T07:43:21.284574Z digest=sha256:b1481d0ea3de1fc6215352eeee59c4333d798bf99767eb60c7a01a6402fc7d4d

Observation 92d6a179-4695-4e99-92a4-c596e43b8601 · inbound

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents cites this paper.

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:56:55.521293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T02:45:25.694378Z digest=sha256:a958648a4f0a11e84fd69431d68e9812c54faafea1b341ec95ff96d7f21ef2ad

Observation 842d94d0-100c-48b9-8371-829985f2993b · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.067368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:0a64fceac8ff960c13158fc288654253408511b170004a13111a2d3edd83cdaf

Observation 08ff4005-a93c-4e45-97dc-15cb1ad52b5a · inbound

DRIFT: Refining Instruction Data via On-Policy Data Attribution cites this paper.

DRIFT: Refining Instruction Data via On-Policy Data Attribution A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:18:54.722039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T01:57:05.784589Z digest=sha256:d09cec9a550744dbd3aa9baf77b9e1a9bd22a11afc2942ff1d1221456e725e71

Observation 7d607ce7-cda3-4fb0-9d51-f107ba4b1af5 · inbound

What are Key Factors for Updates in RL for LLM Reasoning? cites this paper.

What are Key Factors for Updates in RL for LLM Reasoning? A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.575477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:015424268eb7f5628a72dd7917af0e9328ec4e43e5f9cc3ed8e0fe24aacc79a8

Observation 6299d2f4-a2e8-4d50-99c9-9b112124ab68 · inbound

RL Post-Training Builds Compositional Reasoning Strategies cites this paper.

RL Post-Training Builds Compositional Reasoning Strategies A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-09T04:05:55.457832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-09T04:04:13.942267Z digest=sha256:ef77ea6679d20640670b7aeb003822c0cf07bea28aebd34f4713eb69609412f5

Observation 9d906ae0-a2f9-433e-821b-42cf4472e9ba · inbound

Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity cites this paper.

Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T04:32:37.522325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T04:32:37.522325Z digest=sha256:5aa884fb7c832201c7283408bbe28fe8b236917aa4130ed198537bd2f52448e5

Observation 8b0b7d76-777d-494d-a718-80aaea0fcccf · inbound

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches cites this paper.

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T14:45:23.980380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:45:23.980380Z digest=sha256:233a13b5da1997331a7b85c2ac3a321443410dd8e1dd735cc49bfca071bb152f

Observation f9475cb6-7154-4997-af99-64cca06655f9 · inbound

Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning cites this paper.

Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T13:54:23.273919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:54:23.273919Z digest=sha256:cb7da633cbcb0c1879bdcbd4c80f9694cee040aa72218a410b484e2ca27fc1e0