Pith. sign in

Paper Citation Record · LEDGER

ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 37 inbound Pith citation observations for arXiv:2310.10505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.10505 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 37 of 37 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:57.409103Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.661453Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 262fd3c7-4766-406c-81c5-c18a13cbfd8b · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:53:38.919778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:418b7d5204dbfe00ea26b1ed861f9645b7310002a6b5ca6dca8fc7aa2609481a

Observation cacf610b-102d-46f0-8143-b28a42d80033 · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:23:31.123445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:19146f919da500c322e140c670ed44f357ff65bca5413cbc7e796bb37d84ac45

Observation 2e2e35f4-2831-4741-985a-97dad4e1fcc7 · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.409103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.409103Z digest=sha256:ec5a1726d5239803f499bd685b67f6847847da1c3d2a12cc41a1b4fa8a2ea0f0

Observation f0776b47-d595-408d-8199-0e20cbb831ad · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.349648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.349648Z digest=sha256:720b809aa4bcc7e6843f5994efb1a44169bc4d490a62c8e46c6bb01cff9e6738

Observation b3b842df-d41e-4957-991c-f0ea6963c5f2 · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:29.011230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:29.011230Z digest=sha256:5fae6af789e08fd7692a2f19cf5f3f98b87aff4feb08961e373c4110dade651d

Observation d6b7d1d8-ef1f-423c-9de4-6e08705465b3 · inbound

DeepForm: Reasoning Large Language Model for Communication System Formulation cites this paper.

DeepForm: Reasoning Large Language Model for Communication System Formulation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:47.066674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:47.066674Z digest=sha256:8ba160c2a215de6fa6dfe795c92494cb949f7ce00e66512ccdb1c10c9ed52ef3

Observation 8ce80491-f06c-4686-80c5-003108961082 · inbound

Formalizing Learning from Language Feedback with Provable Guarantees cites this paper.

Formalizing Learning from Language Feedback with Provable Guarantees ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:18.715700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:18.715700Z digest=sha256:02d5663c3c3ce28cccea98d90e3ccc00091940cec996fa736528feb81ce7cc64

Observation 6cb29d71-2aa2-4135-8315-01a3eedbb761 · inbound

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy cites this paper.

Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:38.690014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:02:38.690014Z digest=sha256:7a9d846b800ee62c55daf8bcb608d627fce53ed6860a5936d704afd8b3d30e21

Observation bf6ead8d-9d7c-45c8-babe-ddda1bd24138 · inbound

Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training cites this paper.

Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:31:27.344567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:31:27.344567Z digest=sha256:003778672df2b245570ba1eb99bd5c149b26071a59b9e35698438c60f37ea2aa

Observation 62ef0cf6-7c98-4fcd-90de-0686f7a6e8d3 · inbound

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) cites this paper.

Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved) ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:46.084746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:44:46.084746Z digest=sha256:4d9d3b066292f3348f39d92daeb7c3310986fa6c034061142816113757e46d40

Observation c6447613-bc92-4abc-8e7b-a10de149a9dd · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 236

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:07.156651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:07.156651Z digest=sha256:82a5c8e10e02120a9ae9fc529abc5debce5556cdce62860702bbe68beb238420

Observation b485999e-b648-4a9f-9df2-dd6221a6a049 · inbound

QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting cites this paper.

QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T17:39:23.364289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:39:23.364289Z digest=sha256:511a46f8f9f531143e27cc5bfc59db9ad789f086db1c6d4caa09eddf48bc2d79

Observation c232298c-300c-4099-974f-f48cb920172d · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 297

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:02:24.813875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:8621ff116d6b0dad85ab8f0656ff3ffdc5c3fa8f8c19e3a5bd0208f9f275ed2a

Observation 144fb14d-6962-4659-acae-7b7eda5d96b0 · inbound

Inpainting-Guided Policy Optimization for Diffusion Large Language Models cites this paper.

Inpainting-Guided Policy Optimization for Diffusion Large Language Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:46.988775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:46.988775Z digest=sha256:a56b4047212e4d007f366f00f4016d5ed11e62b471bd02e5bd8447a1069335fd

Observation 43b9192e-4db1-478d-89fb-1d29d469c7c2 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T19:38:21.049607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:38:21.049607Z digest=sha256:479285a383eee1624f732cae06c2909272ac588047caab9624a2ac712e9013af

Observation cd5d6fde-b090-4875-99ab-5af041d96c5e · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.815569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.815569Z digest=sha256:882198d6ac4f4af7e10b1c1da95a6753804f0b9a18f4375ad25cadcc5a8fa0a8

Observation 5dbfeab3-23f4-4704-9446-35d3bdb5c8ca · inbound

Image Diffusion Preview with Consistency Solver cites this paper.

Image Diffusion Preview with Consistency Solver ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:43:34.676931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T21:42:24.098676Z digest=sha256:1b5ea0631663388c916c49b6ef8521a8fa5d36b6ba762709237a087341bdce1a

Observation 3f8af7fc-181a-40bd-8b32-1bbb3796e112 · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:49.367579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:47:47.051969Z digest=sha256:4cc88264dc41d12a13d2ad25badbff384f02097409e5d0d8721b3a62eabe4834

Observation 5b776141-93a9-4d6c-83c9-d12f843224da · inbound

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic cites this paper.

Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:39.414831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:39.414831Z digest=sha256:02f66eefcb58deabbe7728cd7f1c9c8038788f193dd025aa5c690cbbb1c07199

Observation 13264c45-0831-4b7d-90fd-ce0557c6a033 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:01:08.766184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:12:06.985609Z digest=sha256:17f20bac0a80a486771463ec30f19631ffd09341d88186b5247c34e8aefeee94

Observation b66ac405-3a0a-4ced-a0a4-2cbd2d320290 · inbound

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation cites this paper.

Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:46:00.902942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:04:12.454268Z digest=sha256:02a2ec94656b9bd214cd262bca5c9264831bea5e69e78d611e3dfb9155582ffe

Observation 0162721a-59df-465d-87c7-96a8042457cc · inbound

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment cites this paper.

BitRL: Reinforcement Learning with 1-bit Quantized Language Models for Resource-Constrained Edge Deployment ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:46:42.470028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T04:23:26.079298Z digest=sha256:527e89a8d377871d741bb78b192b97f7a6e38eeb2bb4693a4749070a6569bba5

Observation a3d2df48-282d-404f-9325-17173f4a6ea6 · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:10.372171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T13:55:22.422923Z digest=sha256:9de82aeb7bbcd325a70b2e8c7cce15a810185823200b9d33ed73b81cd66fbcc2

Observation 8cdce20e-1b2a-48c3-8dca-d2c56cc51c9a · inbound

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex cites this paper.

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:04:04.452971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T09:02:28.407064Z digest=sha256:d3405d67b8b71b4596a1c71b52dc2dfcfcbe622e0fb276e2c1244328ae209cf7

Observation 68130dc2-706a-4799-bdea-c49286b38394 · inbound

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits cites this paper.

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:01:29.028245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:24:03.186413Z digest=sha256:a662d547827798432821c2ece49b34201fc604b99608d7dddbdaac503baa9dc9

Observation a798f6dd-ee22-49f2-83d9-6ee7af1dd054 · inbound

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning cites this paper.

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:22:06.818446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T02:19:27.345348Z digest=sha256:497237db924617cb2db76a2c60d473b0f1e3d8839512f05c7320967d25591df5

Observation ead685cd-698e-49f8-9719-4c5930606f63 · inbound

Self-Supervised On-Policy Distillation for Reasoning Language Models cites this paper.

Self-Supervised On-Policy Distillation for Reasoning Language Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:43:22.208303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T14:42:55.368104Z digest=sha256:9489869cf562740de2494bdbbdf143f183f392a38674864441f452e4d7b37be6

Observation 2f2f5fa2-f0a5-45b5-bd3f-952a6fc66c4d · inbound

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR cites this paper.

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:18:07.233735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T07:14:31.613251Z digest=sha256:fd071c1c7510280ce064953ac46ea82925d07452811e136ac6158ceb1bd1f394

Observation fb06da29-d7a0-4cb1-a23a-27508f0ed17d · inbound

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis cites this paper.

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:09:51.805818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T08:06:02.176934Z digest=sha256:c877235e0a1c5fa94d581869c70208d6b374e81f5defdb9929dbfaa9e2c85b46

Observation 0a1d2f8a-47b6-474f-b7d9-c4d8ea84497a · inbound

Explicit Critic Guidance for Aligning Diffusion Models cites this paper.

Explicit Critic Guidance for Aligning Diffusion Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.787110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:18:46.456767Z digest=sha256:7461ff3815f4d99dfa682407c8ec802aef05b42bbafb3e5f44974c63ef19ccc1

Observation 87373631-0bff-4bec-8c62-2e845eab17ec · inbound

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs cites this paper.

Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:13:26.588012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T12:11:00.402276Z digest=sha256:929333be18ec8d138536bcbed50e22e97b964e0b52b9b848b2f054a906965ea9

Observation 64710dcb-d09d-4c00-99ab-0a71fbe94a6c · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 215

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.678360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:e1b03c92b2f35d919d516cb37b706e4f3f5108508f7c87fc9ac061b1e9da85ae

Observation 3b77b3d6-6413-4afb-be35-9b8cd231d91f · inbound

Rethinking Groups in Critic-Free RLVR cites this paper.

Rethinking Groups in Critic-Free RLVR ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:58:47.902442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T03:22:14.586529Z digest=sha256:98378b564046f5f65db1e54e2bf06dc57dddc4477bc9fb8d98dd6759db481376

Observation 2618604b-2e5e-4b78-b562-3485ea2422f2 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.663139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:55501ba5015f2e9cf39da04bbf9499b0dbc0b5be90e1d3b7a571a011d6a6c37d

Observation 5db4b711-2172-4457-8992-12fff20a36e2 · inbound

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents cites this paper.

RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T14:43:39.668059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T14:43:39.668059Z digest=sha256:713d926c0f15259481a478822188b7af55f116da14596ccdbb111437a982985b

Observation caa5c229-72ee-42f7-9314-b035d470078f · inbound

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion cites this paper.

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T07:14:53.020270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:14:53.020270Z digest=sha256:afd4bfe709f9abafe234a81f036e1e05a2a0deef5ef589b1528496e7c151dbe7

Observation 807dbbc4-4eec-4054-bdb4-cf5469ab7aff · inbound

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models cites this paper.

Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T07:02:36.374494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:02:36.374494Z digest=sha256:4c27334a41c6599a647dbcc57bc9d26cdf37efc78a693d0e881b788e9c34b751