Pith. sign in

Paper Citation Record · LEDGER

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

As of 12 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 5 inbound Pith citation observations for arXiv:2604.11554.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.11554 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:07:52.017037Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:32:02.473074Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T03:05:55.340981Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact16
  • verified fuzzy1
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ae3668c-a270-4107-8c10-7735919be773 · outbound

This paper cites Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:03.310604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:7bbd5ac0928097cb9c63c7128e6e659223e6b439c8c5fdbcb66fba3c460cae4c

Observation 714e39aa-db5f-4aef-aeb5-1c9337d642b9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:16:03.341044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:f2e9609b1098faced5255f30fc8c2b94125eceff18a106ffc5f672e7d6531e10

Observation b274a53d-5f5f-4918-a1b5-a213f762e5d9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Proximal Policy Optimization Algorithms

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:16:03.301588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:fef7dba8db0317852d2fbcce7843784cb50d03c5ed8c69277a494325b3aec73a

Observation db06b6f2-ac58-48de-bc7c-201a510860e9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:16:03.321578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:2bed0197613a3523c6e3764140c0699cdc51b58b7a7afd89cfd8d79ca6b3603f

Observation 389b0b5c-4d69-4e40-ab41-6bacfd400e4b · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:16:03.332879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:8b225091d42ece28bdaf28906301bf7218c56f3564e344f41ff1344b992bbd97

Observation 69cea288-5a44-4e2d-9c81-121c0dbd02d6 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale HybridFlow: A Flexible and Efficient RLHF Framework

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:16:03.315145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:417b89a87967173e382bbbec81b74b05b0b74ad26da51cfa766ce4f6a3b97c44

Observation b83aa4c6-ea2d-4b2d-ab97-d57879b95d3a · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:28:57.414349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:f9db4d05dea2b4feb5d362a0544a7a6b098f220a4d9521aef9d55fb0fdaa66c8

Observation a80d0625-6c83-4771-ae17-5c4bf08115af · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:24:22.137431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:396eed80e9dacd411fa971eb166fa5adc82d3bc292f34330b9421800b03d955a

Observation aa820f69-54c1-4543-b167-7b144dec4266 · outbound

This paper cites AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:03.252155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:7a3edc4154895e93da9556d975737bf68f53dba53ee1e2b4683ad2fa04efae20

Observation 3a2c3f59-b823-4fc0-9825-6df5dccc7e83 · outbound

This paper cites Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:03.246392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:b2f528f7771a822f3274914bdb007b1afd8847906b1559bac511594984fdb67b

Observation 3b3b2430-9bd8-4b90-b944-7c5de00d4679 · outbound

This paper cites type": "function.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale type": "function

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:16:03.270343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:c2c22540f2355abfcd663299291bfbf88944bf83b8161d450cc522a00dc9f537

Observation 591aac5d-eb66-468d-bf3e-096d54507a30 · outbound

This paper cites slime: An llm post-training framework for rl scaling.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale slime: An llm post-training framework for rl scaling

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:51:59.969766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:603d1530009138c541ee6f8d70928afb1c03b6ccced70eecee5387a3ba634cfd

Observation c5e76a10-61fa-4300-8191-63716353c914 · outbound

This paper cites Ray: A Distributed Framework for Emerging AI Applications.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Ray: A Distributed Framework for Emerging AI Applications

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:16:03.296707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:3fc3066169f905b1422df1db8431619ca3102c8228cc618486d91cc8d0db8deb

Observation 323d3a8a-39db-4ade-95b4-8f65f6b6faac · outbound

This paper cites Deep reinforcement learning from human preferences.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Deep reinforcement learning from human preferences

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:39:30.882444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:122e93eefc021ae8e05c9518f5d720b6d25099e076c2943f2177c5960b242c29

Observation b387c5ab-57a5-4f31-8afe-d48cc3ad668d · outbound

This paper cites Training language models to follow instructions with human feedback.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Training language models to follow instructions with human feedback

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:16:03.207345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:379b5b80ccde36fb2620f0ced3a97be426b4aa780abb0a6f37b48b3b89c3d182

Observation aa881d94-6290-4ec3-84c3-a3051ebe5043 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:42.146675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:b352bcbb3b7266de4d6237fff4011c93d8b91839b6db79d2bff2fd81cf850bd2

Observation 8ec5f994-e0ca-49e6-95ef-2cd0bf0003bb · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:16:03.240975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:a846aef5493c6dbc5dfe75eef1c9ec7964b9c11dac81bea077e070d15577034d

Observation e29febe5-7e34-4270-a0c3-f8dbc3898095 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale SGLang: Efficient Execution of Structured Language Model Programs

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:20:01.305558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:74106479bbc96344d0df813a7fe4604ba1a4d63d5a8be7163f7e47b36b3c0472

Observation b2f940ef-fad4-45de-a12e-be7f400cce3d · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:15:20.153254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:68950fc01e92d581fe532e1e31f7dd8c5fb0714819b488c85354cb4b67c5ff5c

Observation 609eb919-2099-427a-962b-6cb59eedc62b · outbound

This paper cites Inference-time scaling for generalist reward modeling.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Inference-time scaling for generalist reward modeling

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:03.235452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:a86a30246374787515a9a8d329c54fb992efffdfa7658700e5e58d373aed026b

Observation 87eebf59-fac5-45cc-b15c-4399c9236ab6 · outbound

This paper cites EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:03.279034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:af121e7d5db8ffc34e3bebf8467587fd30cf77c0e71d6d764c4f49d51576154b

Observation 71c61e6f-ee4f-4fe5-a518-e46df2326e7a · outbound

This paper cites an unresolved cited work.

Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-17T16:51:59.972786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:07:52.017037Z digest=sha256:fcdad2aee044389ded40f350b9d68216e39abd601ddc6182bc6497310b6eb8de

Pith citing papers

Observation f287ec78-98c1-4c7c-a4b5-d6bdf99027d9 · inbound

MinT: Managed Infrastructure for Training and Serving Millions of LLMs cites this paper.

MinT: Managed Infrastructure for Training and Serving Millions of LLMs Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:27:51.877050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T19:25:12.407148Z digest=sha256:27987ec040b57897cee5c5462fb82efb384e80b6ff2502d65be876ac8c97f635

Observation 884a07e3-0ce2-459d-b434-52d7660279fd · inbound

MinT: Managed Infrastructure for Training and Serving Millions of LLMs cites this paper.

MinT: Managed Infrastructure for Training and Serving Millions of LLMs Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:05:06.192120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T21:47:00.295144Z digest=sha256:30830f5e33a1febb8de8010c46c270072a1856eceeb721fd07a74cc5aec15e4d

Observation 6e2c39ff-7e1f-479a-b9c1-8b983d4ab74c · inbound

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence cites this paper.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

Reference 124

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.342846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:60b90919698fcea0155a372cc5f14970f0d3b5120057bf092cbf3cf202577f31

Observation 0e0cbcdd-cab6-4b7a-adf8-2a79aacbd96a · inbound

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models cites this paper.

JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T21:32:02.473074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:32:02.473074Z digest=sha256:07752da73104ca018ee8c577b8faec0f0904be3e89b810f898a2cc1e2e14ce0b

Observation 7b849eeb-c909-49a0-82db-ab62884f9c63 · inbound

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning cites this paper.

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Relax: An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T09:49:31.413822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:49:31.413822Z digest=sha256:cd2388f68e34cd3b09dd1ae7ce3daf37f37501751ef53b6c4402bbc2070d888d