Pith. sign in

Paper Citation Record · LEDGER

ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2502.21231.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.21231 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:41:45.459582Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:27:28.048568Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8017b579-2402-4f6b-95f9-c90fcc2523bd · inbound

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training cites this paper.

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:15:09.582899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T21:12:22.201810Z digest=sha256:7e4d54a2a2009b28f9cea646467dfd47e01a1c7fbd4c8b47658d61a67bd5c51d

Observation d062ed69-fb33-4f4d-9da4-7b47739030ca · inbound

MAGI-1: Autoregressive Video Generation at Scale cites this paper.

MAGI-1: Autoregressive Video Generation at Scale ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:31:15.775808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:31:15.700943Z digest=sha256:c1ca4e78c92633232bea3e46154cf3a60dd2ec0459c4aaba43be0ce1c3809629

Observation 033a95d0-8d15-4b90-991a-67655237a820 · inbound

InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training cites this paper.

InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:06:27.261784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T14:04:31.017142Z digest=sha256:bff488296d0791073b71bf16632822a0c961187dac275efa1e3685a16d056001

Observation a492123d-5b69-401e-b410-6bbfaff96408 · inbound

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training cites this paper.

MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:44:19.608821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T19:44:04.833504Z digest=sha256:44150eeff3af4198a270d198a52cf88b49044a4421d0a56b8ea5e8bda9622864

Observation dc9975fe-da9b-44d6-a3bd-52eb37215dba · inbound

GLM-5: from Vibe Coding to Agentic Engineering cites this paper.

GLM-5: from Vibe Coding to Agentic Engineering ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:46:40.953528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T05:46:40.836161Z digest=sha256:c43ffb56a483111d324af42504cf3ba5c6f67f927cd339f81674922159de1b8c

Observation e5ff70bb-1d1b-4c95-87ef-838c8574ecd0 · inbound

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference cites this paper.

MCAP: Deployment-Time Layer Profiling for Memory-Constrained LLM Inference ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:54:48.395862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:54:33.897112Z digest=sha256:a23a560ee368b1903296aa7fbbbf6736077a1edb91c52b52013927ff6d3b9d45

Observation c1b504b3-a5c5-42c3-8af3-6e14e81c1f74 · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:06:14.982178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:ac83e498107e7c796b81251863f32b10e7d729eeb4581471accd6811cc40a6a1

Observation e25cb165-2ed7-4dd4-b56c-c6e1d6cf8167 · inbound

FlashCP: Load-Balanced Communication-Efficient Context Parallelism for LLM Training cites this paper.

FlashCP: Load-Balanced Communication-Efficient Context Parallelism for LLM Training ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:27:28.050061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T18:10:34.962717Z digest=sha256:ad6d2f1268902a8e3221f2ebf0a4ef330052a5f90d9b5e677005f331370a1170

Observation 4e602654-14e9-497d-ab0b-7d5d49199b30 · inbound

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget cites this paper.

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T00:41:45.459582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:41:45.459582Z digest=sha256:0f5424adc4359b0bece97d55efb77a9ab926b821b07aa4da81b7d0a91fbacce3