Pith. sign in

Paper Citation Record · LEDGER

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging

As of 9 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2502.01804.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01804 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:27:57.561399Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d8853291-31a6-4438-838c-e8177ea277fc · outbound

This paper cites Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.391563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.391563Z digest=sha256:9f6cc2898d7d6c7c90e383ee6449dbefd5710b9378b710f873dff2a5d1faf16c

Observation 36c7597e-646b-4b7d-9bcc-b4aa6344e198 · outbound

This paper cites Ensemble of averages: Improving model selection and boosting performance in domain generalization.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Ensemble of averages: Improving model selection and boosting performance in domain generalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.396007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.396007Z digest=sha256:addc2a7c8696c84f456bc5a602f89759e8913ac854bd89a51cb71ed58ef937de

Observation dd5c57d8-ee63-42bb-95dd-e47857dbbeb0 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging On the Opportunities and Risks of Foundation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.400087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.400087Z digest=sha256:802e520bd8c0b5fda8ea04beebca8370ffe9fd1c42776816fae2ce1c329adc52

Observation f408d395-f482-4108-a0ad-dba645e9922d · outbound

This paper cites an unresolved cited work.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.404425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.404425Z digest=sha256:85c48d27b308b8f645f67c7d2724f4dfe04747618211ed02bb6d17634964e073

Observation feb98c8f-41b0-4247-8458-6caf21c33376 · outbound

This paper cites Fusing finetuned models for better pretraining.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Fusing finetuned models for better pretraining

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.408331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.408331Z digest=sha256:bcf286e0457e79448da28950349584b8d156f80080f7ad8f44ef06b5802d4605

Observation 60b0b9d2-5d8d-402c-a46a-6a79e9e9d94b · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.412285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.412285Z digest=sha256:532470c1dcd5f60ef37a565894747275ab12c27b5ab5e05e9d50627e7fa18234

Observation 73f5ab54-1ee5-447a-8adc-e30fed190887 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.416519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.416519Z digest=sha256:9be5682e743e48300931544750e22b841cdc0c2572ed6c6a05e0852195bfb94c

Observation d5fca2ec-fa38-4895-bcc6-21d032854460 · outbound

This paper cites Pareto manifold learning: Tackling multiple tasks via ensembles of single-task models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Pareto manifold learning: Tackling multiple tasks via ensembles of single-task models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:58.086718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.420385Z digest=sha256:07d8ef7951292c3fe851166e14291fd9804aa9cb08e1542cafda4b8ca7e14883

Observation e84ceb55-0f61-4ce1-b195-7ba48fdcab56 · outbound

This paper cites Understanding Emergent Abilities of Language Models from the Loss Perspective.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.423803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.423803Z digest=sha256:bae383d8976e9a1084e091552bcb746f0c71d5f751f09b68f4c758a24f1c3dec

Observation 78e44852-42e8-403c-b09c-12ae7f0cb577 · outbound

This paper cites The Llama 3 Herd of Models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.427749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.427749Z digest=sha256:eb9d396129e482ca2d0fa19f4626a7d797d814ae7b5d31ab75bcb4f5a0ab3af3

Observation 253c7a5c-2696-4055-82b8-a1b0bab31ecc · outbound

This paper cites DoGE: Domain Reweighting with Generalization Estimation.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging DoGE: Domain Reweighting with Generalization Estimation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.431339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.431339Z digest=sha256:303f268c46717e90071ccd7626c5f985dcd377fd8c1535f88322d543fe961696

Observation 161d34cb-2394-49f7-9871-818fd68db1ec · outbound

This paper cites Dynamic Gradient Alignment for Online Data Mixing.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Dynamic Gradient Alignment for Online Data Mixing

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.434779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.434779Z digest=sha256:ed60286a66b75535bd26f96d1fdea3d7b06b9d9e96decfe0498488f7775e1697

Observation 11e84925-235f-4796-9d8d-0207ad581f70 · outbound

This paper cites A Review of Sparse Expert Models in Deep Learning.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging A Review of Sparse Expert Models in Deep Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.438208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.438208Z digest=sha256:f0bea44a6531350558c4ba81cf871dc6075678b3c6e83fd56c749fe98e049a82

Observation dfe26220-49f0-4b71-a9b3-2ebac4de105f · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:58.071793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.441375Z digest=sha256:7015eaccce2b75007656b27471e23558ffbf3510dfb04a559fdf5c141babe5b6

Observation 51ea79ca-b8ea-4326-9a21-2ceaa9d89f8f · outbound

This paper cites Language models scale reliably with over-training and on downstream tasks.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Language models scale reliably with over-training and on downstream tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.445085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.445085Z digest=sha256:b07b8a9538b66a3d7100a5897b996984fc5a373de8df0f035e13c77af573fea9

Observation e56c246c-31cd-4f83-a061-a65fcaf2b04a · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.449094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.449094Z digest=sha256:69f426d0fc33d79a600137ae82a53303d9b3936c2526abbb8ddee4c1dc1245f9

Observation c269444f-8b4a-4398-a1f9-52d7e5fa46c4 · outbound

This paper cites Demystifying Prompts in Language Models via Perplexity Estimation.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Demystifying Prompts in Language Models via Perplexity Estimation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.452868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.452868Z digest=sha256:7ff82feec905c920a97f6e6b4d877f3fdb6f815afee1425ff950b0c4ef104410

Observation 82923042-9307-4469-b0eb-b2f623bfbf6a · outbound

This paper cites Adaptive training distributions with scalable online bilevel optimization.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Adaptive training distributions with scalable online bilevel optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:58.059125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.456462Z digest=sha256:b0980acfd625294f7855cec3fead228c6927c3947d6362b7c8380e04061f3781

Observation b08e76ba-e5e0-4789-ae2f-033f557d9885 · outbound

This paper cites Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.460128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.460128Z digest=sha256:4d0ea9a4228b4ec09cae5dcd670913400aff9fad560ecacfb86decc381f5aa38

Observation 84c69eca-4bc3-4985-9bb4-c61d6c11baf3 · outbound

This paper cites Hard mixtures of experts for large scale weakly supervised vision.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Hard mixtures of experts for large scale weakly supervised vision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:58.046874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.463977Z digest=sha256:3df2baeb7033c548ce36250c01a81d8ae2af9bd1cef9041f72e3ae6f805792a2

Observation 88c8c2b4-1ca4-47fa-8016-cb09b1ebf898 · outbound

This paper cites MiniLLM : Knowledge distillation of large language models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging MiniLLM : Knowledge distillation of large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:58.034269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.467487Z digest=sha256:87b6b50ba41ddaea222f3aadd9d4d57d5d380e1fa2816c287bf26fc9921f71ab

Observation d14bff1a-2729-4624-a010-25750b77502a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging LoRA: Low-Rank Adaptation of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.471035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.471035Z digest=sha256:52efc94f072e71d4944b1a5313932eab916574d34952121e7161d884fd1f3ebc

Observation b40ae42d-0f3e-4bb2-b432-76a066acc317 · outbound

This paper cites LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.474822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.474822Z digest=sha256:5dffea97b1d1857e9c297f8c6fb4d9de553c802f367109b47ca71a88fcdec10e

Observation 0c026c98-4752-4671-8734-56aaf4aeb769 · outbound

This paper cites Editing Models with Task Arithmetic.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Editing Models with Task Arithmetic

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.478584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.478584Z digest=sha256:7d8234bec186b3c1c199755898e1189d7aa6e5f0499d2a4a005ad34858b546ba

Observation b3d6253a-fe63-4121-ad08-89e0ec34babb · outbound

This paper cites Mistral 7B.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Mistral 7B

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.482212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.482212Z digest=sha256:0f0d432d9869d3cea3306da4430ca187aa0db178d356781e3e07f02ed073b0d9

Observation dfde294a-c936-4492-9067-f943d5e6baa1 · outbound

This paper cites Mixtral of Experts.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Mixtral of Experts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.486070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.486070Z digest=sha256:521713d02344017c244b4150781a6172b19125adec750e6b617fa345d2e85a2b

Observation c19c0c4b-8781-48de-b369-9a39376106a6 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Adam: A Method for Stochastic Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.490125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.490125Z digest=sha256:c828c89fc369863e2910b3130bf6f85b046d4b3e7b680279fd0b2009d850313a

Observation 2745ea1f-7427-405e-851a-79a3be989fa8 · outbound

This paper cites Scaling Laws for Fine-Grained Mixture of Experts.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Scaling Laws for Fine-Grained Mixture of Experts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.493957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.493957Z digest=sha256:db61d947d02b4c02e335be0160b175ed7f8d54763a574ff6e8f692df0985addd

Observation bf38f286-0ecc-4312-a1c4-35e60eafe9e7 · outbound

This paper cites Evaluating quantized large language models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Evaluating quantized large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:58.021236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.497889Z digest=sha256:a0651d227f85fbd832374ff352165abf1869acc2e4e1be51d365463939fdff67

Observation dfa78e05-8bbf-47bd-9645-58cc9e40dd9a · outbound

This paper cites LLM-Pruner : On the structural pruning of large language models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging LLM-Pruner : On the structural pruning of large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:58.007734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.501526Z digest=sha256:b97dabaa6e5fbdf2fc186db1e4cb074b65d05697421080cfdbc02236af6ac0ea

Observation 58b1f7da-f0d7-44b4-91dd-d26fec265c93 · outbound

This paper cites Task arithmetic in the tangent space: Improved editing of pre-trained models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Task arithmetic in the tangent space: Improved editing of pre-trained models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:57.994722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.505793Z digest=sha256:debb0c99520d1c477f77ef729ea4236635449c3c6ae07b0b4bcc8b48d878e706

Observation 427b3fff-99d2-4f59-8d94-36d76c0383d9 · outbound

This paper cites Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.509397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.509397Z digest=sha256:81c42e09d836fa27a19d027acad30732bf029285faac3d297afd6f32b530a847

Observation 4f337655-4d5b-4423-9baa-1ac280c93dfa · outbound

This paper cites Diverse weight averaging for out-of-distribution generalization.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Diverse weight averaging for out-of-distribution generalization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:57.981377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.513283Z digest=sha256:0d6fe3d12a9c8d5ac330357beb7f6c7e5a60014b33c237dfe7c1c16f8480dfc6

Observation 46644d6d-fc48-48b0-bc87-de9253bcb8f6 · outbound

This paper cites Model ratatouille: Recycling diverse models for out-of-distribution generalization.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Model ratatouille: Recycling diverse models for out-of-distribution generalization

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:57.967482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.516804Z digest=sha256:d7d2c85f4a85b16e965e92a39aa8b942c4d67646f36a9c573a9473fda05172f3

Observation 27462a9f-c869-4aaf-9e3c-930437dfe036 · outbound

This paper cites Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:57.956064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.520952Z digest=sha256:dea553bba397bfbd8c9f2f08f9c3e6115dc3e9f027689eaf8ec43775aca1f592

Observation c4d32ccf-b590-4ca8-b125-6fa924230722 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Gemma 2: Improving Open Language Models at a Practical Size

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.524414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.524414Z digest=sha256:6ff23aef96ca574dab1f9060de874536ab9a5142cdca164e647656e2c390571c

Observation e4f29297-d2e5-435b-b527-e50f9296ed6a · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.527885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.527885Z digest=sha256:ef591e52b81abf1979c424f1f0bfb97becb2fba619709cd0602d1e413ac64080

Observation bdf10fbd-1173-4d8d-92af-32373de6e98e · outbound

This paper cites Realistic Evaluation of Model Merging for Compositional Generalization.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Realistic Evaluation of Model Merging for Compositional Generalization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.531857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.531857Z digest=sha256:83182cc4e5a195bd45420446efa4771042dd504624efe8c8cc81934fed172a17

Observation 025f7d6c-2553-4731-87aa-2563046a66ed · outbound

This paper cites Efficient large language models: A survey.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Efficient large language models: A survey

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:57.944220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.535349Z digest=sha256:7ecb554b26872b0eb226ce24e178bc7e2d92f8a3fbf92d266a626cf312d08377

Observation 3043a95d-cf58-44b5-9822-d5c72ff727a3 · outbound

This paper cites T., Wu, T., Song, D., Mittal, P., and Jia, R.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging T., Wu, T., Song, D., Mittal, P., and Jia, R

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:57.931671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.538832Z digest=sha256:d06450773409a3e676cd49a4e33f1c1dcd0e1d638f7ce432e2ea3533f00c12e5

Observation 2faa9b3a-7af2-45e2-becb-87d3f888bdad · outbound

This paper cites RedPajama: an Open Dataset for Training Large Language Models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging RedPajama: an Open Dataset for Training Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.542454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.542454Z digest=sha256:ffcb09610771937c81be19846fafa7598d58d1e195fa69af89dcfd8a528a71a1

Observation 4d0d2575-1cf5-4df2-a85a-ac6c08caa8c3 · outbound

This paper cites Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:27:57.919199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T14:27:57.546212Z digest=sha256:22cb151e02ecfaf323dfbd0bf856e2eecc54c189be615117a88f946605d8e20e

Observation cda5a5a6-bfab-46eb-a2c2-8ed6b42dd7dd · outbound

This paper cites Structured pruning learns compact and accurate models.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Structured pruning learns compact and accurate models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.549919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.549919Z digest=sha256:efe2a28643d15a8d69c821cbd964eae0dff560740a763603bb8a92f23f2f6f04

Observation bb2ab740-b7eb-4d1f-ba23-f443ee314bdf · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.553441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.553441Z digest=sha256:f2b2a0e200c94c04cabcf21fece121e3fdf1765020880fb7075902f1aa07d4a9

Observation aa434cf1-ebd4-4aef-9ead-0802752e0920 · outbound

This paper cites DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.557385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.557385Z digest=sha256:0ec08bfa9654773d3430615acc57675a778c6c49e1eb40048b940240c57af673

Observation c1156ac8-7ef6-4773-9ce7-1b3eacfd4c53 · outbound

This paper cites write newline.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging write newline

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.561399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.561399Z digest=sha256:0d2645a0f665d29c2e8eb14b3054dae01014aa8ea6bb887e0ceb333033deb107

Pith citing papers

No inbound Pith citation observations are available.