Pith. sign in

Paper Citation Record · LEDGER

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 1 inbound Pith citation observation for arXiv:2605.08639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.08639 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T01:38:03.841465Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T08:35:16.435272Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T12:58:08.745243Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact24
  • verified fuzzy54
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dff35b60-e754-4dac-a87f-5f0fbb4a0890 · outbound

This paper cites Outrageously large neural net- works: The sparsely-gated mixture-of-experts layer.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Outrageously large neural net- works: The sparsely-gated mixture-of-experts layer

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.999101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:6a9ea7bf1aa9f055bde6dd2c2e7d215d80090a34c71186e54ded9d4b49a61b66

Observation 52b4106b-93ce-4111-9933-d6c751b7684a · outbound

This paper cites Gshard: Scaling giant models with conditional computation and auto- matic sharding.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Gshard: Scaling giant models with conditional computation and auto- matic sharding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:59.024843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:ebdcc64a1714b7a7bab8d52233c18e7615fd21a10921587e31a7b36bee64132e

Observation f179413e-f69c-4a37-9723-4df3e915bebe · outbound

This paper cites Switch transform- ers: Scaling to trillion parameter models with simple and efficient sparsity.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Switch transform- ers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:59.016199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:2f3c9db9cfaf17eb85bcec690ead913f44072102ca12ee8b88976e725df967e7

Observation ecd82fc7-d778-4c60-91dd-5f0394840717 · outbound

This paper cites Introducing DBRX: A New State-of-the-Art Open LLM.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Introducing DBRX: A New State-of-the-Art Open LLM

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:59.004084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:ca973d3f8c42ee5ac2045913b3baddb01e74af6393f3e791363548a23c34f0a7

Observation 46044026-8da8-4b38-bebe-02303127681c · outbound

This paper cites Mixtral of Experts.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Mixtral of Experts

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:14.870115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:d69d4ae56f9541198e34555618f283d6f37f34441f8c261a45ce7c2fb99d29c1

Observation 85e6d9d5-5143-4400-88ae-828e652f1cf6 · outbound

This paper cites DeepSeek-V3 Technical Report.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning DeepSeek-V3 Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:14.867213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:e5b8c089a28bb9c41fddb76ee52deb85efde524acaf98f2465c996395ba7a9bf

Observation af418330-8429-4356-97e2-f18e217f2f0c · outbound

This paper cites Qwen3 Technical Report.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Qwen3 Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:14.876080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:73643005fdd871eae11dfe583dc576930fee312c970d67a536f71b0509ceee15

Observation ef79a0e6-6687-4f80-9749-e6e0136ae0d7 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Kimi K2: Open Agentic Intelligence

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:14.873230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:2811bc862d3c8851dc35e1de1d3a0e95b6c108aa8a2459fbbafcd26b00b077d9

Observation b9beb139-041a-48c5-8ca1-e0ea276e2ab6 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:84846a617df6042d321263db09b7382cc5fc3f7944d2262a14ab960be85d56cc

Observation 4ff822f2-3ccd-4f9e-bc5b-7bb03601c371 · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of- experts.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Glam: Efficient scaling of language models with mixture-of- experts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:59.020714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:aef49a8b3646abe77c62ecc801714d727298dd67effbf23d03ce50a8b32f8b70

Observation 21a61d10-097b-4f92-9226-519e5e265d26 · outbound

This paper cites Deepspeed- moe: Advancing mixture-of-experts inference and train- ing to power next-generation ai scale.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Deepspeed- moe: Advancing mixture-of-experts inference and train- ing to power next-generation ai scale

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.994171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:f1f4f23ba9da8a6100b347ffadf76aa7cdbccbdec09615b97a4d6a5dc635211a

Observation 568f501b-5448-4eeb-8040-19bb1e4d07bd · outbound

This paper cites Training language models to follow instructions with human feedback.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Training language models to follow instructions with human feedback

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:59.011896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:44105c14150e8e0d749865db27ae7fcf06f68d86b84870d8c2bea9886cf93b1e

Observation 71d13990-1ed4-4027-9c1e-577e84dad624 · outbound

This paper cites Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Deepseek-r1 incentivizes reasoning in llms through reinforcement learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.989243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:7d2eb124033adf8b7c3a7afbd61fd26d2958cffc49529a6c793d649c3df694c7

Observation 77bbf96c-60d6-4a09-8f5b-296fff3f7779 · outbound

This paper cites Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Flexmoe: Scaling large-scale sparse pre-trained model training via dynamic device placement

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.963005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:45244d20fab8deaacf2bbae11dcde47309b28b5da3c9fa5715ab802405fe44ec

Observation deb4980a-5017-43e3-991b-5a59f9b5f5e7 · outbound

This paper cites {SmartMoE}: Efficiently training {Sparsely- Activated} models through combining offline and online parallelization.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning {SmartMoE}: Efficiently training {Sparsely- Activated} models through combining offline and online parallelization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:59.008240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:b8428156338e6751c7bafb1da95cbca1583c9896774e4a78b0509536811e5262

Observation 7b2f7151-8195-4cec-9a5e-1b88bcb32ff8 · outbound

This paper cites Micromoe: Fine- grained load balancing for mixture-of-experts with token scheduling.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Micromoe: Fine- grained load balancing for mixture-of-experts with token scheduling

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.773337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:14e7d6da0f54df8129ad67ea3250902fb2a5c5adfb16ac3a020aa194ae19a7b5

Observation f9ad5cd0-4e1a-4ecf-b654-6c62915a71ab · outbound

This paper cites {PopFetcher}: Towards acceler- ated {Mixture-of-Experts} training via popularity based {Expert-Wise}prefetch.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning {PopFetcher}: Towards acceler- ated {Mixture-of-Experts} training via popularity based {Expert-Wise}prefetch

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.884605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:73f1444b30aedef39c9d2d696f5bd81295201aa2fc35f156c4abeae05123a45d

Observation 5bd5b82c-3c4c-4369-a810-19e5545c25f0 · outbound

This paper cites Laer-moe: Load-adaptive expert re-layout for efficient mixture-of-experts training.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Laer-moe: Load-adaptive expert re-layout for efficient mixture-of-experts training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.935971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:3370af32e706fd24ea7f2885b9fec0d29d92465625df2fc9f0d232f59f8a5492

Observation 385433df-5e23-49b7-8fc0-dbd1ef95d4af · outbound

This paper cites Moe parallel folding: Heterogeneous parallelism mappings for efficient large-scale moe model training with megatron core.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Moe parallel folding: Heterogeneous parallelism mappings for efficient large-scale moe model training with megatron core

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.788840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:62fe0a839468ed5b86204962ee551f4225642e63e26c1b16cc48e5d52582c71d

Observation ce6e0d2e-58a2-4e5f-a990-bdaa7bf10f97 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:14.863718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:90d1801830691659866fd22e17ec34ad61608fec40431a0e6d9054577a082ae1

Observation e6dc883e-635b-4b41-9755-5044ab1dc1b7 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:14.776908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:ab5f86d5b2bf42c9984b479f4b1c4625dbab51ac1333124d2bb30059d663f5e8

Observation 1b77baf5-dc1f-4d01-8e6a-f27350964c5e · outbound

This paper cites Competition-level code generation with alphacode.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Competition-level code generation with alphacode

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.794293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:603c79a7d4fa8038997963ef8e8002c19d0eb6674ebb3fab545464e68c8169a8

Observation 1be9df3d-fb67-4348-aab2-2fbd64496764 · outbound

This paper cites Generalizing Verifiable Instruction Following.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Generalizing Verifiable Instruction Following

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.784748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:ad56aa89f09dea8c697a2fec4b1f957b6a837d88fae88a95d1759fef6b977116

Observation 2c3bc7b3-94fa-49e3-a0a5-3fbd26ae892d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:14.852010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:4d635ee544108d547a663e2864951611059739c426b31d5e198c40cca0048d1a

Observation 5d5d2efa-4075-4991-9e0b-cf73ed260b87 · outbound

This paper cites Expert Parallelism Load Balancer (EPLB).

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Expert Parallelism Load Balancer (EPLB)

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.880483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:5b4f194bfcc6bc9cf4a72b36ec3132c184c4299267a40a34490c0addedcbe47e

Observation 624d296b-b2b6-4098-be53-8a4a9e9998ae · outbound

This paper cites Human-level control through deep reinforcement learning.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Human-level control through deep reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.931027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:c2b92604e483e80b3022aaa81e2ab6f19e7f11367baf7b787564eefb663610a4

Observation 857db13a-d183-4cc8-8bb7-a5ce7a4eaac1 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep rein- forcement learning with a stochastic actor.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep rein- forcement learning with a stochastic actor

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.940727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:09d4c4e5b68a5e2455521a6759936e3078a882a883869f4ba4385f44d6ce7cc9

Observation dc74dbbc-b1de-43a9-bfa9-aa9bd083e354 · outbound

This paper cites Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.805376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:5fdebfbd31091ba378301abc3089dbe6fe7511beaf7be8b731221c47ff0ab6fb

Observation a6f72534-1f13-44d6-9a98-16638ee235e6 · outbound

This paper cites Stabilizing reinforcement learning with llms: Formulation and practices.arXiv preprint arXiv:2512.01374, 2025a.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Stabilizing reinforcement learning with llms: Formulation and practices.arXiv preprint arXiv:2512.01374, 2025a

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.833281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:ea96984de6c92fda75a4e55303999e2d13247261b2173a89b60356d1ee8942f2

Observation f2f9a668-a7c0-4f77-b354-b157fbe567a3 · outbound

This paper cites Optimization and approximation in determinis- tic sequencing and scheduling: a survey.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Optimization and approximation in determinis- tic sequencing and scheduling: a survey

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.945201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:c2e1cad18a572a9f93ddabd4d3aa7e4601353386a872603b6e994154bc86a7e2

Observation 8e7af441-3f0e-408b-99fb-28eb637f96a2 · outbound

This paper cites At- tention is all you need.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning At- tention is all you need

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.807726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:8411303f41f37c24634c27313e3d7d2c38dc7fdc06477c6f58545fc7622d5b2a

Observation 18f9c5e4-fdf6-46e1-adec-e442c670b004 · outbound

This paper cites Zero: Memory optimizations toward training trillion parame- ter models.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Zero: Memory optimizations toward training trillion parame- ter models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.798827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:df064a30c0266e04e5498f56b40fa9dd3b57f868721e759578e48237cfec1856

Observation 2fd505dd-cf8f-468e-a56b-0ff75cc59c7c · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:14.855673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:46d8e1268b5a16599c781833768ea999e15ef55a233308d731970125ec2dbe8e

Observation 06914383-ba02-41d5-938b-202de464d0a3 · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Gpipe: Efficient training of giant neural networks using pipeline parallelism

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.917601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:1acb36103cc0c738310a6ebfae1b8e91c942f41887bec21a37ae8cab1011c9a2

Observation 32c4b5cc-8779-46e3-b680-9b2c93e941dd · outbound

This paper cites Pipedream: Generalized pipeline parallelism for dnn training.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Pipedream: Generalized pipeline parallelism for dnn training

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.926894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:8e2ffdf63fa08c3f5b6aa8f78f7305d3098626337f4cab33e7be28b603bb38eb

Observation bb3e96d9-2d93-4f65-9896-eca7ea46162c · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Gpqa: A graduate-level google-proof q&a benchmark

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.912608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:833bf17188828586aeb91b07c5bde98ae413fde9ccbeb30e4c3ada89bb28bafc

Observation 65345a14-10bf-4d85-9117-c210ec1ff531 · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based} generative models.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Orca: A distributed serving system for {Transformer-Based} generative models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.903645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:e57f57e1dafaca2cee681a22b7e182e4100e97fb75eb1fbef80efdcab9992a4b

Observation 9317bfce-1349-4e73-8f1d-47f89af2593f · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Fast Distributed Inference Serving for Large Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T11:23:46.844514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:2e0585e84cfcf96383877aea1bc56b3f059d8c62c843be46faac0073925896f4

Observation 9fb964f2-6b02-4965-8b77-0340e6bd5fd0 · outbound

This paper cites Proximal Policy Optimization Algorithms.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:14.829234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:e16d5a87278b96f30dc305a56ce587bb9fd585da2976a674c95f8aefb2fc9850

Observation e3303144-1b90-4337-984c-f1e09be065ac · outbound

This paper cites Taming the long-tail: Efficient reasoning rl training with adaptive drafter.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Taming the long-tail: Efficient reasoning rl training with adaptive drafter

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.908296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:06a49e4cbd4a0a99f4137a392ea98419c7320f330fe91d53c013cf62e535528e

Observation 7806e4fe-deeb-45ae-a94f-b608b56dc170 · outbound

This paper cites Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:14.780484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:c92bbb4916528b0a81157f04746b6e5024364f2ca0168ab27096f319e9913e70

Observation d543da68-bab7-4ee2-b074-5d912626231c · outbound

This paper cites Optimizing {RLHF} training for large language models with stage fusion.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Optimizing {RLHF} training for large language models with stage fusion

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.922164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:77babf44e339ba5cc5fada3a886b53fee85f644282be99fe18280d95643e6560

Observation b66f0bbe-6517-44e0-8c0b-9ed262dbc682 · outbound

This paper cites Towards efficient reward service for rlvr with request- level flexibility and batch-level constraint.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Towards efficient reward service for rlvr with request- level flexibility and batch-level constraint

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.958891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:5b477cc2fe99a1bea68edb59ef0056d6cfb258b08ee9173de678b7e97ef9f9fa

Observation 3aa6a3ef-4103-4653-b148-2dc98ca7e855 · outbound

This paper cites On the uncapacitated location problem.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning On the uncapacitated location problem

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.894780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:4afbedce4e38ae2d65f93fbdfd07b07ae26f8618299b3806c9c65d6e7598941e

Observation 96e454c5-a57f-4fcc-a685-22268ca5fdfb · outbound

This paper cites Linear-Programming-Based Load Balancer (LPLB).

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Linear-Programming-Based Load Balancer (LPLB)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.975665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:e6d28bf3765a4c13bac4f7dd46ad8cdc2674a3f1270958f915b9fd0184beacc0

Observation 20129cc4-2d8f-4744-9263-5d1853b3153c · outbound

This paper cites Alibaba hpn: A data center network for large language model training.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Alibaba hpn: A data center network for large language model training

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.980058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:6862a2bf47e4a8e04eba26841bb272191c2ef277cecca15355c545ebe543a48b

Observation f36e49a2-b4ff-4c96-b2c6-285808a35645 · outbound

This paper cites NVIDIA GTC: Accelerating Mixture of Experts Train- ing With Rail-Optimized InfiniBand Networking in Cru- soe Cloud.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning NVIDIA GTC: Accelerating Mixture of Experts Train- ing With Rail-Optimized InfiniBand Networking in Cru- soe Cloud

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.890336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:22bd2bd2e21a0b4dd0953c21fbd98cec68e650c14d51aef0e2e05d4f187a2d23

Observation 34f396d0-c592-4030-93ba-9fda397948d6 · outbound

This paper cites GPUDirect RDMA.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning GPUDirect RDMA

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.898809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:6028a6fde1b48c36c320cc3c5a7ecc8f9fa6a2c8455f0bdfc4b825e7e5ed66db

Observation 68807d47-bed7-4b07-9a64-eee85f945339 · outbound

This paper cites Optimized primitives for inter-GPU communication.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Optimized primitives for inter-GPU communication

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.966717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:ce26bcf098ebb516b6ac38233926cc1d537e6366e8d3a227e198db13e9e2f21c

Observation 1b5027e2-0064-412c-9107-87ab95e2bbd8 · outbound

This paper cites DeepEP: an efficient expert-parallel communication library.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning DeepEP: an efficient expert-parallel communication library

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.984310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:c63dddd1abc59a6c0ec9acdbb035cda7cc083ebbceb83f0f6e57d66592a8443c

Observation 1dcc35f2-baff-4af2-8ad4-820eee02c33b · outbound

This paper cites {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning {DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.866727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:7514c7b4d08322e551bebfe17e1d38019425e2bc8ddebfaff4fba786dbbdef08

Observation 28703dcc-1496-442d-935c-d775a0033ecf · outbound

This paper cites Megascale- infer: Efficient mixture-of-experts model serving with disaggregated expert parallelism.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Megascale- infer: Efficient mixture-of-experts model serving with disaggregated expert parallelism

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.859126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:fb87c854fb1666452bd84650d3ab0c26fd3e639af38da13d7f04391b402ad0cf

Observation b8c3b01f-b5d5-4568-a3f5-9f383d52cefd · outbound

This paper cites Disttrain: Addressing model and data heterogeneity with disaggregated train- ing for multimodal large language models.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Disttrain: Addressing model and data heterogeneity with disaggregated train- ing for multimodal large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.876273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:a0567f79991518d0184f4a08842dd95c6bbce85871aa9a739a208166c40d7864

Observation 0d51a2e2-d65a-4ec5-9a50-15321224479a · outbound

This paper cites Heddle: A distributed orches- tration system for agentic rl rollout.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Heddle: A distributed orches- tration system for agentic rl rollout

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.769379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:26c9625cee392e7be32a255e6517d15a76f45e2a065654c348c2c0cfa77982ed

Observation d4f8504c-7e28-43c0-9149-7e01a9ccae01 · outbound

This paper cites Bounds on multiprocessing timing anomalies.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Bounds on multiprocessing timing anomalies

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.871762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:c3d77defaeb92cc8fa1f3f6c9314acf361afa1532f474e98a381a086706ffc6b

Observation c5a7f206-30f3-45e5-bddc-ea9dca27d731 · outbound

This paper cites The SCIP Optimization Suite 9.0.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning The SCIP Optimization Suite 9.0

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:14.793302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:9fbe6d80c938bf448fedf95902be04bc5bbda4e87aae741b78bf76279906af79

Observation d0431f56-14c5-43fa-89b8-55732931a853 · outbound

This paper cites Slime: An LLM post-training framework for RL Scal- ing.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Slime: An LLM post-training framework for RL Scal- ing

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.852858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:3d07e4a0e6f2f38b68337ba75456ddceb81b0cb61b05776e4b1a76d598204bc0

Observation 2be4ef82-d728-4b9f-8c3e-65e47d51bf21 · outbound

This paper cites GPU optimized techniques for training transformer models at-scale.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning GPU optimized techniques for training transformer models at-scale

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.785432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:4896736292a0a1ec74a2093c87de0eabafd85744706e3b2f362d89e260cb6cab

Observation 3af36f7b-c481-4563-8992-d00f68bc5b66 · outbound

This paper cites Sglang: Efficient execution of structured lan- guage model programs.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Sglang: Efficient execution of structured lan- guage model programs

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.843832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:f6b1a37d3749400a2a7f9806b008b6a9db0e85db8ca49c05ac0928c3437e1ef3

Observation 2a85d91b-0d66-485c-861e-6ef7454e3f95 · outbound

This paper cites Open Multi-Processing.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Open Multi-Processing

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.839050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:8072a7621535f174f2cdc82d3aa3d2d13033468f7b1cd0cc94b7043d83897957

Observation e1477a93-dc66-4203-b21a-10b453d2a826 · outbound

This paper cites NVIDIA Hopper Architecture In-Depth.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning NVIDIA Hopper Architecture In-Depth

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.848374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:9f34415be91a02505eaacb7feaae3f93b66ba80f573e51e6caa2c072723ee6ea

Observation 3f851cad-76ab-4529-89ed-ed56d4e247e6 · outbound

This paper cites an unresolved cited work.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-05-14T06:28:58.862623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:dbbc36f199305a1cb0567947811ce15f76c8c9b25076498a1407164cfab42c7f

Observation 2b740753-657e-479f-a5a0-3935b7d11cff · outbound

This paper cites Flashattention-3: Fast and accurate atten- tion with asynchrony and low-precision.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Flashattention-3: Fast and accurate atten- tion with asynchrony and low-precision

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.830705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:193ad871b13642d3317096a9bb6f2a7103dea94d89084e6221f8c69b187bcf0c

Observation 6832e81d-59a6-4c45-a709-76af756d5447 · outbound

This paper cites https: //github.com/NVIDIA/TransformerEngine.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning https: //github.com/NVIDIA/TransformerEngine

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.835254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:8566e066fd5ffc1c4488cf1e10f9e683aa825a5b32411a8c0dde6fb32625280b

Observation 11be90ab-bfc4-4ab6-8d3d-685362a04275 · outbound

This paper cites Ray: A distributed framework for emerging{AI} applications.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Ray: A distributed framework for emerging{AI} applications

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.855524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:74295276c9681a878241c245dd49df97e51533f5d043c0008aa3723ada38abdc

Observation e7153983-0537-4ef1-8764-059d129ae801 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:14.840788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:bbe4ec195969b1f091e573099e814ecc8841841ee1dc249963c5fffe97647d85

Observation 032e3e72-ffe2-46ba-b2e0-b817a2dac756 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-12T01:46:14.801130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:d1810107c0bf8616ffc6cbf36b06adc3d19d84192c1ad1564a71e5a9bc6d3f3e

Observation a4915de7-acd6-45b8-b6a4-9a4118b9a47a · outbound

This paper cites StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.859743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:e47c952786a609e05af98fcf063a5d50f76a7bbd50bc420fa24da556b81ce174

Observation 80069897-0e6f-4418-aff2-9137cbe7b3c6 · outbound

This paper cites Accelerat- ing distributed {MoE} training and inference with lina.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Accelerat- ing distributed {MoE} training and inference with lina

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.971149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:9be09b7586179cbb1804aac2b834e56faf645232b8ea74e299e44dbde2b3e083

Observation 9a22491d-3788-4847-84dc-16b2a95b2f19 · outbound

This paper cites Tutel: Adap- tive mixture-of-experts at scale.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Tutel: Adap- tive mixture-of-experts at scale

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.825534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:65682005ff322340d88d4880b7b8a4d250bcdf6ea6e6fd829862f06244e4288e

Observation f41717dd-3c21-4f74-aba6-9b0211cc8618 · outbound

This paper cites HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:14.809929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:07e90b0a27c44e284af15c9bee68b599ba88de69f41405777baa6bcd79420119

Observation b4645b66-b82b-4d01-b3da-bd583d60d13a · outbound

This paper cites Megascale-moe: Large-scale communication-efficient training of mixture-of-experts models in production.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Megascale-moe: Large-scale communication-efficient training of mixture-of-experts models in production

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.847449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:48ba8aa77e62da4ced327f4bf8fdcd525dc7f1536c041c3606678b365c3ed401

Observation a00867cb-86da-4ede-8776-33825a23d0d4 · outbound

This paper cites Centauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Centauri: Enabling efficient scheduling for communication-computation overlap in large model training via communication partitioning

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.821197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:5aa280ca35bea06da215924f6339389791efb9316a28ba033c8f8fcc093ff173

Observation 71b27ec6-5411-444b-b686-3bb3b9c6fbfc · outbound

This paper cites {MegaScale}: Scaling large language model training to more than 10,000{GPUs}.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning {MegaScale}: Scaling large language model training to more than 10,000{GPUs}

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.790159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:4ff4108c3e499d0e938986e58d98ee50994e6c7f0e5144067e8bdbf52c4780d7

Observation a999f442-c95e-4305-9e99-7f454c60e9ff · outbound

This paper cites Comet: Fine-grained computation-communication overlapping for mixture-of-experts.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Comet: Fine-grained computation-communication overlapping for mixture-of-experts

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.803518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:544172856e52b056db43c216451e2f406e1a9bcad58de50d6dbc7ac8b9b27309

Observation 77edfa72-05a8-4614-b0ab-497e1a131a5b · outbound

This paper cites FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.822291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:424696394b6a5314253394a419489544fac26e019307013f7778ba2542a2470e

Observation 817bc4fb-959e-498c-8853-31be95e918e2 · outbound

This paper cites Janus: A unified distributed training framework for sparse mixture-of- experts models.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Janus: A unified distributed training framework for sparse mixture-of- experts models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.816731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:699d03440d37e165da1a80cf320609e7814960044ee5e14efffe179bb5c9719b

Observation 276bd38c-7d29-4d78-94fb-87cd95832341 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Hybridflow: A flexible and efficient rlhf framework

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.812129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:14372073b22ebdfd171d8de0f1b0a7362a000b15fb8696281640f5462e4cd9e0

Observation a862306a-ed53-4214-90f2-ca0336c60cce · outbound

This paper cites Rollart: Scaling agentic rl training via disaggregated infrastructure.arXiv preprint arXiv:2512.22560.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Rollart: Scaling agentic rl training via disaggregated infrastructure.arXiv preprint arXiv:2512.22560

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.818155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:1ab9b68d1ef3b444e57ef20868287a701ce378c207a43a29b34a47bd2752972a

Observation 8166be66-f00f-4fe1-81e2-72925456e3e9 · outbound

This paper cites Fast infer- ence from transformers via speculative decoding.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Fast infer- ence from transformers via speculative decoding

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.949381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:80a59e4a060340f6742929cbc931703b8ca2b9c6253ee9946f195d629d6df096

Observation 10568927-0499-45ab-82f8-6c55c919ff32 · outbound

This paper cites Ef- ficient rl for llms with dynamic and online speculative decoding.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning Ef- ficient rl for llms with dynamic and online speculative decoding

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T06:28:58.953806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:ac7cc3942c098e3bc8a88273e5efaff274d15457219c8e58eeeeab8c3874d876

Pith citing papers

Observation b2a876fa-2975-4906-84e9-4d18f3f99ad3 · inbound

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training cites this paper.

Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T12:58:08.746533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T08:35:16.435272Z digest=sha256:31b2a8b1e967f4d5ecd554c04ae1e9af003851f401e88884aae90ef2e6ac3c18