Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T14:27:57.561399Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2502.01804.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T14:27:57.561399Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d8853291-31a6-4438-838c-e8177ea277fc · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36c7597e-646b-4b7d-9bcc-b4aa6344e198 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Ensemble of averages: Improving model selection and boosting performance in domain generalization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd5c57d8-ee63-42bb-95dd-e47857dbbeb0 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging On the Opportunities and Risks of Foundation Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f408d395-f482-4108-a0ad-dba645e9922d · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feb98c8f-41b0-4247-8458-6caf21c33376 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Fusing finetuned models for better pretraining
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60b0b9d2-5d8d-402c-a46a-6a79e9e9d94b · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73f5ab54-1ee5-447a-8adc-e30fed190887 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5fca2ec-fa38-4895-bcc6-21d032854460 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Pareto manifold learning: Tackling multiple tasks via ensembles of single-task models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e84ceb55-0f61-4ce1-b195-7ba48fdcab56 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Understanding Emergent Abilities of Language Models from the Loss Perspective
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78e44852-42e8-403c-b09c-12ae7f0cb577 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 253c7a5c-2696-4055-82b8-a1b0bab31ecc · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging DoGE: Domain Reweighting with Generalization Estimation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 161d34cb-2394-49f7-9871-818fd68db1ec · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Dynamic Gradient Alignment for Online Data Mixing
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11e84925-235f-4796-9d8d-0207ad581f70 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging A Review of Sparse Expert Models in Deep Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfe26220-49f0-4b71-a9b3-2ebac4de105f · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 51ea79ca-b8ea-4326-9a21-2ceaa9d89f8f · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Language models scale reliably with over-training and on downstream tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e56c246c-31cd-4f83-a061-a65fcaf2b04a · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c269444f-8b4a-4398-a1f9-52d7e5fa46c4 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Demystifying Prompts in Language Models via Perplexity Estimation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82923042-9307-4469-b0eb-b2f623bfbf6a · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Adaptive training distributions with scalable online bilevel optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b08e76ba-e5e0-4789-ae2f-033f557d9885 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84c69eca-4bc3-4985-9bb4-c61d6c11baf3 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Hard mixtures of experts for large scale weakly supervised vision
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 88c8c2b4-1ca4-47fa-8016-cb09b1ebf898 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging MiniLLM : Knowledge distillation of large language models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d14bff1a-2729-4624-a010-25750b77502a · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging LoRA: Low-Rank Adaptation of Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b40ae42d-0f3e-4bb2-b432-76a066acc317 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c026c98-4752-4671-8734-56aaf4aeb769 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Editing Models with Task Arithmetic
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3d6253a-fe63-4121-ad08-89e0ec34babb · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Mistral 7B
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfde294a-c936-4492-9067-f943d5e6baa1 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Mixtral of Experts
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c19c0c4b-8781-48de-b369-9a39376106a6 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Adam: A Method for Stochastic Optimization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2745ea1f-7427-405e-851a-79a3be989fa8 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Scaling Laws for Fine-Grained Mixture of Experts
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf38f286-0ecc-4312-a1c4-35e60eafe9e7 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Evaluating quantized large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dfa78e05-8bbf-47bd-9645-58cc9e40dd9a · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging LLM-Pruner : On the structural pruning of large language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 58b1f7da-f0d7-44b4-91dd-d26fec265c93 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Task arithmetic in the tangent space: Improved editing of pre-trained models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 427b3fff-99d2-4f59-8d94-36d76c0383d9 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f337655-4d5b-4423-9baa-1ac280c93dfa · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Diverse weight averaging for out-of-distribution generalization
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 46644d6d-fc48-48b0-bc87-de9253bcb8f6 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Model ratatouille: Recycling diverse models for out-of-distribution generalization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 27462a9f-c869-4aaf-9e3c-930437dfe036 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c4d32ccf-b590-4ca8-b125-6fa924230722 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Gemma 2: Improving Open Language Models at a Practical Size
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f29297-d2e5-435b-b527-e50f9296ed6a · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdf10fbd-1173-4d8d-92af-32373de6e98e · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Realistic Evaluation of Model Merging for Compositional Generalization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 025f7d6c-2553-4731-87aa-2563046a66ed · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Efficient large language models: A survey
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3043a95d-cf58-44b5-9822-d5c72ff727a3 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging T., Wu, T., Song, D., Mittal, P., and Jia, R
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2faa9b3a-7af2-45e2-becb-87d3f888bdad · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging RedPajama: an Open Dataset for Training Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d0d2575-1cf5-4df2-a85a-ac6c08caa8c3 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cda5a5a6-bfab-46eb-a2c2-8ed6b42dd7dd · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Structured pruning learns compact and accurate models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb2ab740-b7eb-4d1f-ba23-f443ee314bdf · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa434cf1-ebd4-4aef-9ead-0802752e0920 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1156ac8-7ef6-4773-9ce7-1b3eacfd4c53 · outbound
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging write newline
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.