Pith. sign in

Paper Citation Record · LEDGER

Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2208.03306.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2208.03306 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:41:22.247867Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

26
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 93a51336-5131-4fde-a1f8-d295bbeafd7f · inbound

Editing Models with Task Arithmetic cites this paper.

Editing Models with Task Arithmetic Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:09:13.060337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T08:09:12.716163Z digest=sha256:c74d0098a7dc2bced60ba7a1573a3941b55bb778ff5bd20ac965b29dfc6f527d

Observation 1a67f19a-4f70-4387-8c2c-84c82cd66125 · inbound

Task Prompt Vectors: Effective Initialization through Multi-Task Soft-Prompt Transfer cites this paper.

Task Prompt Vectors: Effective Initialization through Multi-Task Soft-Prompt Transfer Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:13:30.125991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T22:11:59.206066Z digest=sha256:a8ff945af57f7f268a8e30012b8f63e095963b96e32969a8ac7dff1302ec7af2

Observation 01bb3cee-ca4e-496c-a52b-66f6bff67e97 · inbound

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities cites this paper.

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:16:04.906658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T22:16:04.386706Z digest=sha256:804afec370882b3b50ba462ec438f3abeb26f5aa99eeb53233786feead9e37d7

Observation 204ed7b1-527f-4820-beff-4f0ec9fed988 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:00.952905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:ac12c499d43122367f9d4c31e1d10ce5b0383f0cec4a80b7de08d9735a4ea9a5

Observation 7b058177-784b-435a-a3ff-7acf76772da6 · inbound

Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging cites this paper.

Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:21.516966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:21.516966Z digest=sha256:3e2ad8aee09faff49d0e881ccf5c54e6f8bd34fbc3701b10f2088a5138142c61

Observation b4b1e656-90ed-4aec-8020-cb5f269abe2c · inbound

NoLoCo: No-all-reduce Low Communication Training Method for Large Models cites this paper.

NoLoCo: No-all-reduce Low Communication Training Method for Large Models Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:02.441768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:02.441768Z digest=sha256:b37e00b94b8adcc4a0d60bffc42b036fc9ec40976ba4bf6c1d51eda3b7199614

Observation c4151b25-40eb-487f-81f6-58055bca8106 · inbound

Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model cites this paper.

Group then Scale: Dynamic Mixture-of-Experts Multilingual Language Model Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:21.433548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:21.433548Z digest=sha256:eba052c3a6559349c313ddae6e76be6d9974f1516c3f087b7ffbb82dc9061cc8

Observation 3719c411-dbc3-4766-89de-cea2be10bd24 · inbound

Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts cites this paper.

Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:56:26.216747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:56:26.216747Z digest=sha256:949a7d90bcd0c9202dfe35cd590252cd49720df3c668f21f40679914b6e4e1eb

Observation 24e3a7d0-fe10-400a-8c73-3d34e3d9eb7e · inbound

CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics cites this paper.

CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:59:04.662169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:59:04.662169Z digest=sha256:f3fbf4c3db88f3184b04b70597dfa9716d4f860adb51190d18db168ba2ea6f16

Observation c829d2a0-d80c-476b-b417-ce674ea619e4 · inbound

FlexOlmo: Open Language Models for Flexible Data Use cites this paper.

FlexOlmo: Open Language Models for Flexible Data Use Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:16.051548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:16.051548Z digest=sha256:3325774b30dd3f2801466d2e00f52c585f97b7d3b6fc6d487bd2585d49f449c9

Observation 8f5852f7-d3e2-406d-bf99-b68c77247d1c · inbound

MOMO: Mars Orbital Model Foundation Model for Mars Orbital Applications cites this paper.

MOMO: Mars Orbital Model Foundation Model for Mars Orbital Applications Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:48:15.103016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T20:47:51.509798Z digest=sha256:94696a46e024beb31060304d00ed415476df5dcad7a0b46ad6aff5d12c7f36c1

Observation 27729a76-23bc-495a-aab5-fcad434f7204 · inbound

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts cites this paper.

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:56:11.347027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T05:52:28.822723Z digest=sha256:960287091215019154a77f8e806fb7c40baf2b4c7abe96b75fb3e7093494eff0

Observation 4602cde6-ba2d-4709-b06d-03129ada670e · inbound

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data cites this paper.

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 162

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:57:21.795138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T05:56:38.042978Z digest=sha256:543c682fe2d683df1d64d3820b7c3e0a444ed0351652f18bbbd9b5acd3d578e1

Observation 3dd25be9-893f-46ef-be1f-1dafa8d7d156 · inbound

HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts cites this paper.

HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:55:04.758253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T05:54:32.496951Z digest=sha256:df030ec7d88bdd38cc870946cfe8beea1c513cdb804cca8dddaf8b6bcb139aec

Observation 30bd7b97-324e-431d-9077-40337fc42078 · inbound

MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts Unification cites this paper.

MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts Unification Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:23:31.881288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:21:10.641070Z digest=sha256:38e4386ae46d625dab66797905470a47c9b85992e4d11df32f676dab47b5d477

Observation 6bef364a-0d99-4d45-9728-c51e1b4e1419 · inbound

ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models cites this paper.

ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:23:16.769512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T12:22:30.263086Z digest=sha256:bd8cc7e4dcbe16861598e67d4dad55c09562368aa4cf41f0aad16cd47506ae87

Observation a6aac792-d8ae-41bb-b993-f2616357ffba · inbound

Decentralised AI Training and Inference with BlockTrain cites this paper.

Decentralised AI Training and Inference with BlockTrain Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:40:00.431608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T23:32:43.585170Z digest=sha256:0095ccee6087049a558a8a0659434e0fd3895e99987481c8fcb8c90a175807e8

Observation cd94ee50-d4a0-4886-b0fe-3a2337f9ecdb · inbound

Decentralised AI Training and Inference with BlockTrain cites this paper.

Decentralised AI Training and Inference with BlockTrain Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T12:29:35.403453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T12:29:35.403453Z digest=sha256:be4fd4a4d898a9efa887f34891682008f45429609655d84a45234fd80a9203a4

Observation 256c947c-611d-4a69-baa4-1ab1b55861fc · inbound

CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield cites this paper.

CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:15:44.043684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T05:46:40.510955Z digest=sha256:c78529a6a07589880138cb414288985a035a7c4c21c35efbebf3975df53f8777

Observation 93570348-03da-426a-ba90-993117dbef4a · inbound

CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield cites this paper.

CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T09:24:14.388774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:24:14.388774Z digest=sha256:6295e0f809431ad2be8ec3aecce95a238b7678cac1438f67fd0d9b860b272ae0

Observation f25ca2e0-bf00-4380-9012-68d56b699d0d · inbound

Identifying Latent Concepts and Structures for Generalized Category Discovery cites this paper.

Identifying Latent Concepts and Structures for Generalized Category Discovery Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T14:47:03.350280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-02T14:42:01.334822Z digest=sha256:4162183304f62a6ba002315a9b6280f4966d470aa82b55688340f245d8c4cd58

Observation f601a8d1-a5d2-4f76-965d-e75617e19f10 · inbound

Modular Foundation Models for Time-Series Perception in Digital Twins cites this paper.

Modular Foundation Models for Time-Series Perception in Digital Twins Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-12T01:22:51.284207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:22:51.284207Z digest=sha256:3fd8896cb24cb68aff2ccb40bb91113bbc46c7fa304ad70d56d44904e86c159c

Observation c2433688-1c4c-4442-a014-c6bea26815e8 · inbound

SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization cites this paper.

SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-08T00:41:22.247867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:41:22.247867Z digest=sha256:4043de6f3beb44370a6c7ea3f1bec70522f59a9aa6ab8ab63544b4dc971d655f