Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2504.11343.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T11:38:53.871009Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T04:05:55.456269Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation f3cfd3a9-54be-4a41-b964-576ef5d6a952 · inbound
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 55773bad-37ca-41b8-a1eb-2acc1865b435 · inbound
Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bb2d36c3-c7c9-4a0d-9b02-8641cb2cc399 · inbound
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c7116e4-8cdc-42c2-9271-48c094d9c972 · inbound
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3a903c09-9f27-48b5-996a-8758fa183abe · inbound
DiffusionNFT: Online Diffusion Reinforcement with Forward Process A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9bb53df9-d8d6-4975-bdc3-ab161ec4c708 · inbound
Simple Policy Gradients for Reasoning with Diffusion Language Models A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d40d08e-cbcd-4cba-a805-8b25870ed109 · inbound
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2e87cb3-0435-455e-8590-bda005b55f76 · inbound
TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acc3680d-d292-4990-a7da-29ed62b2cacd · inbound
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 73b77fc9-59fc-469e-91af-c9034311ff2f · inbound
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1755dd5f-1fb3-48ea-8ae5-caa2491a1599 · inbound
Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 630449dd-b80d-4aa3-a288-bfdc3eeaca58 · inbound
Kernelized Advantage Estimation: From Nonparametric Statistics to LLM Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4e3f625f-f911-47dc-913c-9b7283007d79 · inbound
Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 27f16019-f5ab-482c-a3c3-229879c8f308 · inbound
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8904c37e-bae8-43f4-b1f4-e6b1b9576a75 · inbound
Gradient Extrapolation-Based Policy Optimization A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ac7f4f7a-08a9-4ae7-9c10-ac45239fb51b · inbound
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cffefd4d-5467-42c3-b666-e5c677216762 · inbound
Selective Off-Policy Reference Tuning with Plan Guidance A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04ef7fec-55a2-469b-b11e-13f85b769347 · inbound
Selective Off-Policy Reference Tuning with Plan Guidance A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6d45c1f7-88b7-4b72-9a5e-75ef342eb4fc · inbound
Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3cd6cc3-8198-42f8-b0c1-5d701390201b · inbound
Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fd58e815-33b8-4c2b-a799-7d7328ae7e9f · inbound
Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0aee1525-3fe3-446e-a08f-0a4820de4161 · inbound
ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c20fbf8a-6285-40a2-b777-36c0e8059a36 · inbound
Traj-Evolve: A Self-Evolving Multi-Agent System for Patient Trajectory Modeling in Lung Cancer Early Detection A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 247f0921-085c-40d3-bb1a-5832dc6274cc · inbound
Rollout-Level Advantage-Prioritized Experience Replay for GRPO A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 92d6a179-4695-4e99-92a4-c596e43b8601 · inbound
AsyncWebRL: Efficient Multi-Step RL for Visual Web Agents A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 842d94d0-100c-48b9-8371-829985f2993b · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 137
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 08ff4005-a93c-4e45-97dc-15cb1ad52b5a · inbound
DRIFT: Refining Instruction Data via On-Policy Data Attribution A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d607ce7-cda3-4fb0-9d51-f107ba4b1af5 · inbound
What are Key Factors for Updates in RL for LLM Reasoning? A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6299d2f4-a2e8-4d50-99c9-9b112124ab68 · inbound
RL Post-Training Builds Compositional Reasoning Strategies A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d906ae0-a2f9-433e-821b-42cf4472e9ba · inbound
Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b0b7d76-777d-494d-a718-80aaea0fcccf · inbound
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9475cb6-7154-4997-af99-64cca06655f9 · inbound
Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.