Pith. sign in

Paper Citation Record · LEDGER

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training

As of 24 July 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2604.26687.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.26687 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-07T11:25:24.079444Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-24T06:31:00.690269+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact20
  • verified fuzzy36
  • unresolved0
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch10

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ce38fdb5-67d6-490d-8838-a3eed637cf53 · outbound

This paper cites Adaptive Preconditioners Trigger Loss Spikes in Adam.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Adaptive Preconditioners Trigger Loss Spikes in Adam

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:02.571868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:87fc83bd9ceaba6ad1a3e0b59da6c2dc6e250982309ea7a0540dbfb989deadde

Observation b8151a1e-e8da-4712-a83f-24603cb0ab24 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.193602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:6baf77f747d388d40a63e531f913a376169ba20dcd482ed571c87544293b02a9

Observation 96a41743-27f2-456d-8637-7534efa04ed1 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.278555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:477f5b8d1bd150602e17692303d8e23d2f3e469cd55a9106a51fbdbe4f9f6d1f

Observation fcbb17d8-1fb8-4e8c-831b-20bd2358824f · outbound

This paper cites AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-09T03:55:07.651495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:6a82b9c1957a0d25b6f26ebacec8c08fd81bb0cc86f578bcbff559c469783805

Observation 2f7f0642-a84e-49c0-b1ee-5ecdee47d09a · outbound

This paper cites Hé- naff.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Hé- naff

Reference 5

Resolution
verified exact
doi, observed 2026-05-09T03:55:07.741005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:e12517d64994e2bee5ce5db129e4fde86ec619c2e1d4cc0754180dac56e865a9

Observation ad1c541a-78f1-4a4c-b527-f853e15e41a6 · outbound

This paper cites Irreducible Curriculum for Language Model Pretraining.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Irreducible Curriculum for Language Model Pretraining

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-09T03:55:07.703229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:f273ded5eedab4a49ea5e2ae149404ecde9702452e79749b5eb23627736a3ea7

Observation 1aa72cfd-dbcf-4d09-a66e-2abe74b42c54 · outbound

This paper cites Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-09T03:55:07.782058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:a0c10971fec9b1325a822028ae1324251e6d949ed78b67f4c66bb9119b21b2e7

Observation 8231002a-53cf-4de0-9edc-2c0ccace7da9 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T03:55:07.728713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:51085ea74349d22b80b4d3ea4b22ce0021dee20f02eccfabc743b255c4b6320b

Observation 62284fe0-7eee-49cb-b9ab-f074887264fd · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:20:31.604247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:2dbc5f1123902e9876ee288ccebbc177398b1bc21be37954e2fbb67e1ae89f04

Observation 601ea0b9-43b9-46fd-a632-e1ee13d2f3b1 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.204297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:2869d5d589e22c913cca261cb73c28cf96d7983408320418f414aaa5383934d1

Observation bd7a5c7e-6a6b-4398-874d-1a159a02bb94 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 11

Resolution
verified exact
doi, observed 2026-05-09T03:55:07.653949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:5c6f5ca1d0d46a3fe0344bf6f01be5c63b1ed979dfcd29317ce7d843adf5db69

Observation 9ed7c8cb-a4ad-42e2-b003-3e278f172253 · outbound

This paper cites OLM o: Accelerating the science of language models.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training OLM o: Accelerating the science of language models

Reference 12

Resolution
verified exact
doi, observed 2026-05-09T03:55:07.785088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:03df97b69cdda8027563e67808dcb9776829f7eef0eb8748f113e7182c48549b

Observation 652722ba-7a5b-4793-bef1-fb2294cc0307 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.246526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:f7b5d4123ed85977f2ea2e8964b4a529634c147c403fc53164b5e34bc90846d8

Observation 094cb3e9-1cb1-4103-8d66-6e39a30a3ba5 · outbound

This paper cites Johnson, Pulkit Agrawal, Haijie Gu, and Carlos Guestrin.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Johnson, Pulkit Agrawal, Haijie Gu, and Carlos Guestrin

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.170634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:4db7d1789f2c5d2d90595aceb72cb8bc0fdbb4cf7a0004e1bba525dc1bfe39f3

Observation 42311570-c6c6-4dcc-a2b7-f67cde179995 · outbound

This paper cites K2-V2: A 360-open, reasoning-enhanced LLM.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training K2-V2: A 360-open, reasoning-enhanced LLM

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-09T03:55:07.818944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:ece973d4253d843e19134aa02ec1b9869f424af5998535254deb43df5f02d74e

Observation ec1ff916-7ca9-4e3f-bdbe-59097e3f915c · outbound

This paper cites Scaling Laws for Neural Language Models.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Scaling Laws for Neural Language Models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-09T03:55:07.714200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:a0f7b3a513917bde6a74977ded14285e1647f078a029fe6a93c2aae9365ae34b

Observation 13b29926-4b86-4416-bb98-69a14df19eca · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-09T03:55:07.752948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:4fe9aca2c845145ae23f7d9deaf3375e3bfe190ff3f15f0bed14751bf3d13b55

Observation 35e3f6af-612e-4efe-b773-e499fead9acf · outbound

This paper cites AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training AdAdaGrad: Adaptive Batch Size Schemes for Adaptive Gradient Methods

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-09T03:55:07.660974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:43ee0d148e7bd430a903122264caf4873af7d27da5bd3b0cfa9da96cd6b86ea7

Observation c22be6b6-4e45-4f36-9c61-8d76820745b9 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-09T03:55:07.665859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:d7af860677a74e0c4b3d15d90d2a63c9c15bac7dbeefa5364f62f8f1a09792e4

Observation dfbf74cd-450f-44aa-9d1c-8863fb2a7219 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 20

Resolution
verified exact
doi, observed 2026-05-09T03:55:07.787636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:5eafcb1af70697dc597bd7e667e255d26a357a63172f639d885fc7823a09186e

Observation adf083ba-5f0b-4fe5-9019-adedf364fa74 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.173202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:605086676027cdf4904687b3fdc4f86644fbb22d7b664a77fe38b3ac5a312556

Observation ee3c6446-778b-403b-aedb-a8e4b21c58c5 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.179047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:42bb1e3baa2e31dab4b84582d84014ff807e7fe692aebfb5a7b8441ba7d51194

Observation 8c54e1cb-da70-4bb7-b2d9-f6cc39af5844 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.258468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:fe3daa812e26d2b1fce5a6c28497f380553dbb89d883a23930b78bc56d0e9391

Observation 448f98bd-ef6e-4c1b-a247-d0c76e71e8cb · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.175937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:e600d606546a01d23f871f583cb06e77e34138df89862eab98bbc75fd661bc73

Observation 4504aa58-5da4-4538-a447-c549edd90be4 · outbound

This paper cites Hwang, Luca Soldaini, Akshita Bhagia, Jiacheng Liu, Dirk Groen- eveld, Oyvind Tafjord, Noah A.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Hwang, Luca Soldaini, Akshita Bhagia, Jiacheng Liu, Dirk Groen- eveld, Oyvind Tafjord, Noah A

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.207787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:b4d1159217cfcd3bedd6a8d67419aa8aac2b8ce5b24cc42c62720f0be386e4ef

Observation 7c210070-74c6-43ce-b212-583b1a422cc5 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.273102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:2276dc278f4d527a7511c405615161bed1f77c62e74ee6907a097cc62c014634

Observation 1898ec2f-7c16-4df7-bc05-7ab12df833ab · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.261487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:f0b1576a9600486a6dfb14d5ebb6bbaef17bddd7fc874ef367c0d4ab9d2162aa

Observation f1a18a6e-f1ff-4991-9b06-7a01e79878fb · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.264820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:11c7f9aaf6c5b97afa88232fd757a49897fcaabcc3e20adfb910591c737cca35

Observation 57059ce8-db1f-4aba-bfc9-a47f49585db8 · outbound

This paper cites An Empirical Model of Large-Batch Training.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training An Empirical Model of Large-Batch Training

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T03:55:07.909709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:db166ad307cedbd130c08b0142e4df965c7639450015988a2f3a3e3f9d3b4810

Observation a353efde-e779-4263-b2e9-c4b061dbc102 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.241037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:3c1c1d953901525974d4ee6a71a87f0f4e759cfc7257bafdffdb1e38429c4009

Observation 38df66b6-61d8-430a-ab03-eded9d64606c · outbound

This paper cites In5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training In5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.249364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:51632add5d67f4824663a6ed57d07c8b18750b621122197a1d46d499de1a4e40

Observation 62d6381a-72c1-4f3c-a94a-1598383c7d95 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.220323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:13ea9e8cfe8fc6c85377d8468260d726d2cc567c12b347a277d287811bcee41a

Observation a4247abb-69f1-4074-9ffc-7d73d9262932 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.222981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:eb3c0f04430779c932b07f95de708d4e403f1df81463efa39d3267974457b459

Observation 88289efa-0657-4c64-b8ed-48a7b124cf08 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-09T03:55:07.800411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:0508962bd2df3f7c7c5b232a925850e289bd9bb44f07497c84c994be4afb1b3d

Observation a462a08f-967b-46e3-b1b1-d877fb6e0a0e · outbound

This paper cites Devanur, Gregory R.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Devanur, Gregory R

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:21:25.537054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:58cf778ecb5f104c25336f8d842ea57c86c18087e85ee25853f53f73b6cb3320

Observation 4af3a16f-feb7-40fa-b6fe-0e1d708aec5e · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-09T03:55:07.860344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:fa42df647fe37715a71c2dbfe5789d1ac400e757d4f184e514552d25f19cf0e5

Observation 2bff3644-c8c2-4bf1-92e8-4fe7facd7194 · outbound

This paper cites Ostroukhov, A.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Ostroukhov, A

Reference 37

Resolution
verified exact
doi, observed 2026-05-09T03:55:07.827914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:fead41bf635d498191b0e61d90c9f83a81e6f316845901f4b2372e64214d5240

Observation 60f363e4-5414-4338-863d-ec9fdf5279c1 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.217558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:2486338d54a731876406c1bf389aa2fdeb901f7cacaf212073f4249e7e1b552b

Observation 0e494fdd-c1e0-47b3-ae00-043a44807060 · outbound

This paper cites Ganger, and Eric P.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Ganger, and Eric P

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.214569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:45c77122b2b008d2850ea4ae83dae28259c21866188f60be7a797a7d677b9798

Observation e28580ac-427c-48d3-b5b1-36e5cb95f5dd · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.232460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:1d3f3f6251d0d5adf8177ecbcce03351ff5432672e8ac465864a09c9e4348296

Observation 9579db7f-2923-4297-9f68-ebb539621ac9 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.211408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:a2d053d7afcba0ae0885adbc6af0d71e4836b67d9e4314d81ce14aae41e65e8d

Observation 0658506d-9c48-4f4b-875c-4ac4f0b9d9ac · outbound

This paper cites Generalized Slow Roll for Tensors.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Generalized Slow Roll for Tensors

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T03:55:07.806803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:277c60fa5b5fa5807d95ca7b636e851ad4e0ea918d0ae49aa53af96a3840686f

Observation 4a6664c9-e81b-4e02-865b-7018e7cc4e46 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.226485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:020c41e6589845478077fd8495ace2c852dfa6b460b671c68215fccc9de38127

Observation a1bdddd4-090e-4948-868d-90fb5ab295ba · outbound

This paper cites InKDD ’20: The 26th ACM SIGKDD Conference on Knowledge Discovery and Data Min- ing, Virtual Event, CA, USA, August 23-27, 2020, Rajesh Gupta, Yan Liu, Jiliang Tang, and B.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training InKDD ’20: The 26th ACM SIGKDD Conference on Knowledge Discovery and Data Min- ing, Virtual Event, CA, USA, August 23-27, 2020, Rajesh Gupta, Yan Liu, Jiliang Tang, and B

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-09T03:55:07.812216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:99e68eba07d5b97811cc32ce6d19ca17d64811a99fc94b02e7679352291ff74d

Observation efc689bb-b15e-4727-ab44-20a712fefcd6 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T03:55:07.895409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:44a6ce1a684556ea00758da4d545e0f91da226fbe8d8079537e0fabcb6f4785d

Observation 97b7dfb7-82f8-4552-9540-c1bef589f5d5 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:34:44.898910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:16c8c6e606395d612c2e2cd7b4b268b6b4defddfac9b577aa02ccfed4e971c06

Observation 3e648084-6253-435f-a355-63e530943292 · outbound

This paper cites Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.235470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:eb75cefdbd4067e90b1b766e9a111fdf45d6e2889a8f3a23c67ebda181c40733

Observation c7733b6a-8bcb-4450-8ec0-f7a95808bf52 · outbound

This paper cites In6th International Conference on Learning Representations, ICLR 2018, Van- couver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training In6th International Conference on Learning Representations, ICLR 2018, Van- couver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.229753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:08428fd431734d6830a3169faf99b52a5dda9f8ab459ba56b25e499c5ca35409

Observation f89d5de7-e57f-40df-9c8d-865baaf4814a · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.252282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:027eddc77e6253f224869f7d1fc00a8aee5e2d92ac8a83355ff1ae7dd3692ac1

Observation e2d80675-c1b7-4bd2-978d-de83bf8685e1 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Gemini: A Family of Highly Capable Multimodal Models

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T03:55:07.749487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:5765347057007579a4532c2d3ba06ef98fd69bdbec5311018a29b6b89431e5dd

Observation 6753e2cf-983b-46a7-b1e8-de90ce7ba1f5 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-05-09T03:55:07.833943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:19bbfcb9400fe499928e8f60aede9f5f57f72d30274314b0463cd0f633d39e03

Observation 6526f3a3-3686-48fc-81c6-552acecd1825 · outbound

This paper cites The Llama 3 Herd of Models.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training The Llama 3 Herd of Models

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-09T03:55:07.645264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:dcca6c7fae91f24aef222b70f6548e74f7d12d8383d86c4d304d39ac1feffb17

Observation 6bb2cda6-9ac3-4662-8c68-5c5959482551 · outbound

This paper cites Qwen2.5 Technical Report.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Qwen2.5 Technical Report

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:21:25.540607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:ec63d1d8e0e4f2594fd8641bab92f6c46d7e17cd68bee5776354a45ddb0a343c

Observation f9f512a6-56a9-43b5-84b3-d0703a05d5a0 · outbound

This paper cites Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment

Reference 54

Resolution
malformed identifier
doi, observed 2026-05-09T03:55:07.697092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:334e5216130f1bcb230e009f4e5b95828c837a6480def21713b8cba75c22f231

Observation b6d9384a-fe9a-4a2f-96fc-3ae36ac4a281 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.276074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:f777a97106bc1dbd68d89b206b3c62ccb944edc00bde9715620f5cf14832b402

Observation 15484cfe-048c-4c9d-9cdd-bbb80a3e2661 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.198936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:b39fb63e1e9db3c6f0e910935571fbdf5649d6ae2ff8f410342dcd982e6dadeb

Observation 13b3fc24-bf70-402b-aaf6-91fd9aca806e · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.196263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:7dafc793d91550db5b803771f1338e01e1d2d689c619b1e0b13f91f231683d65

Observation 74e294b7-1deb-44d3-bce2-d53319c5316b · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.201466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:f745da87aa64256e2f719ba0b0f88f713ec96cfdd8d7d6e58cbc51eef0b24714

Observation 4194a9f5-110f-4367-b78a-85e22cb590c8 · outbound

This paper cites Le, Tengyu Ma, and Adams Wei Yu.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Le, Tengyu Ma, and Adams Wei Yu

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.255331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:fa60a0f1c03f41d32acdad2631776c5d524ad421253a78f6d951c83802283cde

Observation 70e483b8-ae0c-4069-8124-f3b5418d3260 · outbound

This paper cites GSPMD: General and Scalable Parallelization for ML Computation Graphs.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training GSPMD: General and Scalable Parallelization for ML Computation Graphs

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T12:36:36.562059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:6f1554c035b63e83ed2c26597cefc794ef706c48c9067f70aa60ffe8263af10a

Observation 4649e375-7598-4900-a0fe-95109555b71b · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.190890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:f064b1a97bad6cdd584bc5c592676d28f3cbc7a253efde6a773bef8d1eb5c68d

Observation 96145f5e-085a-467c-9748-aaffa710ea47 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training OPT: Open Pre-trained Transformer Language Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:53:19.162556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:b629958b7ede01dfe90fe3d71ea6a6a12c0a35cb476a0ff5d5d9dbc3f574bb83

Observation e2c47dff-c6eb-4775-99d5-821c32b01db5 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.243780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:1cb812f639a25e97ffe5a0b4f48a9c92c8c886bf46acbfc757a777384b09bf1f

Observation fcc0b277-6dc4-4f22-9536-8ea38b11997a · outbound

This paper cites Xing, Joseph E.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Xing, Joseph E

Reference 64

Resolution
malformed identifier
raw_fallback, observed 2026-05-27T07:08:46.267897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:0ea0986b847f7b39bca44fe5516a8da615bd03f9f51272358350b8048f99a7b1

Observation 78c1d1ae-2c4d-45ad-bf0c-a3210fa3b712 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.270669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:4e7ca9e8bb0bccd105ac9c679656dab234522ab5934233ab2dee38982bab52d1

Observation 86d65717-f9e5-422b-b5f8-c6c962e63744 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.187749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:56603e738ecfd629541959baeaa23ba9c9d9215b554908d05604b8b5cb6f207c

Observation 0a6ce1e6-3564-4032-8c83-afda1d06ee0c · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.182078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:d484707e66fa8563624aacee7cb9f987045b07a15c73b0f63fee171035384c4e

Observation 6d90268e-13e9-469c-adbf-b1e422569307 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T07:08:46.184935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:c7728084b49601f19baabd81f7bc0fccb34e7f5e815f5532859bc4631dfedfe9

Observation 1a2cccab-2ebe-412e-882a-4da2052ee7b6 · outbound

This paper cites an unresolved cited work.

COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training Unresolved cited work

Reference 69

Resolution
malformed identifier
raw_fallback, observed 2026-05-27T07:08:46.238088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-05-07T11:25:24.079444Z digest=sha256:c8ec1afcb8f66eff6ebf2971baeac4b7690c95749040f48b6d90c5224b0c8b40

Pith citing papers

No inbound Pith citation observations are available.