Pith. sign in

Paper Citation Record · LEDGER

Why Do More Experts Fail? A Theoretical Analysis of Model Merging

As of 18 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 5 inbound Pith citation observations for arXiv:2505.21226.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21226 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:44:36.927598Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T02:42:45.039852Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T11:24:38.144959Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e29d4997-1560-4b85-9b1a-a27a5b968a5e · outbound

This paper cites Evolutionary optimization of model merging recipes.Nature Machine Intelligence, pages 1–10, 2025.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Evolutionary optimization of model merging recipes.Nature Machine Intelligence, pages 1–10, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:32.608750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:32.608750Z digest=sha256:27f5757aaf59fb5557ffe5e4e854e9908319dbcba90a76ec160e4e1ef3f8a802

Observation 9a59c3fa-eed0-4c63-975b-6ee3caa393bc · outbound

This paper cites Living on the edge: Phase transitions in convex programs with random data.Information and Inference: A Journal of the IMA, 3(3):224–294, 2014.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Living on the edge: Phase transitions in convex programs with random data.Information and Inference: A Journal of the IMA, 3(3):224–294, 2014

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:44:38.665865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:44:32.677745Z digest=sha256:3fdfd28630afb104abf7455d655e742909c49ab961e6b155f1f77b9a0c2a193b

Observation 5b7ccb7a-ad41-4bfe-bc71-3eea40db3453 · outbound

This paper cites Program Synthesis with Large Language Models.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:32.790393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:32.790393Z digest=sha256:b88c6bda565b3e46f694d5fd665914ee2d60c51133e19f6e7a5873cedbc5a221

Observation d265a83f-900c-46d6-ba8d-a30de0a616bc · outbound

This paper cites Revisiting Weight Averaging for Model Merging.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Revisiting Weight Averaging for Model Merging

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:32.908553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:32.908553Z digest=sha256:c88c705ef8fadd72a337e4fa08ae91573fd3f9d4938781590ae5217b14fe2d26

Observation 733253d7-b3d9-40ee-b755-673c64c1964c · outbound

This paper cites Model Swarms: Collaborative Search to Adapt LLM Experts via Swarm Intelligence.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Model Swarms: Collaborative Search to Adapt LLM Experts via Swarm Intelligence

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.011741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.011741Z digest=sha256:a1aa2889e2669aa710a94040a0fac0732a022e6cdba8d43f8c3740e338d8eaf7

Observation ffdab5c8-5420-4a83-8fa6-a02dcf3303ea · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Measuring Massive Multitask Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.131628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.131628Z digest=sha256:4f4bab71f83032989236d3d3f93abd6da40b2e34db7de5325ce71901c2574953

Observation d5213ade-951a-4a9c-aa43-dddaed611e6b · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Measuring Mathematical Problem Solving With the MATH Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.275322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.275322Z digest=sha256:0ad6210501d37ee4e13b1304c7ffa1d4dccbb5508d78b7550e6c7658f0606717

Observation 7de861e7-b0a6-4fa2-b594-db4a575157f0 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1 (2):3, 2022.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Lora: Low-rank adaptation of large language models.ICLR, 1 (2):3, 2022

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.398276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.398276Z digest=sha256:98d726fa5817e32089de554d42dbee20888572dede2173f77ada6c9df371304c

Observation f0422f54-d67a-40e6-af3e-8cd41f517f1e · outbound

This paper cites LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.469848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.469848Z digest=sha256:0bcf68f2333c033e9023a27e6503a7e0075c965ba4f6de55ed059b3f0172c44c

Observation 9ca7eda8-fd42-4db0-af8b-0cf7fd2f754f · outbound

This paper cites Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.633023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.633023Z digest=sha256:5eee60512cbcbc7858704ae9ea7e08fe0158c76b8337de3b7603f77830ddb6ab

Observation 2088241f-d960-4be2-a8f4-5cdd46cdb00d · outbound

This paper cites The singular value decomposition: Its computation and some applications.IEEE Transactions on automatic control, 25(2):164–176, 1980.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging The singular value decomposition: Its computation and some applications.IEEE Transactions on automatic control, 25(2):164–176, 1980

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.713425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.713425Z digest=sha256:fb3590cfdab224797c678c916c6c80505cd92b16a2e961d4643038874a34efea

Observation 45931957-3294-4814-9ff7-98b9280b7feb · outbound

This paper cites How many degrees of freedom do we need to train deep networks: a loss landscape perspective.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging How many degrees of freedom do we need to train deep networks: a loss landscape perspective

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:44:37.342378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:44:33.847753Z digest=sha256:5910efc3a072ef0b7187060783b11a1fea7534a550cc3726c57f56f0045ad449

Observation 44f2716b-30b3-49e4-b820-8e7757dcc472 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.957436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.957436Z digest=sha256:0c7af0c540b02b12509380385ae39dc74b7d6c50b781dd0d2b0a6feae30fe22b

Observation 7eb2afd5-079c-4563-99bd-828a43b0c9fd · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:34.061461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:34.061461Z digest=sha256:8319629acb67839dd6f44bf12ce691b91bf457c51f05ea56c64da8de46a9083a

Observation 7c15703c-d038-4a43-a856-03f26bf24f44 · outbound

This paper cites Gpt understands, too.AI Open, 5:208–215, 2024.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Gpt understands, too.AI Open, 5:208–215, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:34.182206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:34.182206Z digest=sha256:fd8b1c226b02d81dbdc737800a5139b6c4faec96da379c730544845759107fe6

Observation 1c215081-916c-4983-9e2e-78cf062b1f82 · outbound

This paper cites HiFT: A Hierarchical Full Parameter Fine-Tuning Strategy.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging HiFT: A Hierarchical Full Parameter Fine-Tuning Strategy

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:34.329800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:34.329800Z digest=sha256:6894f182159424e7b92b3e30e62c3e138cbf8aab62556ad262b2067700fffe74

Observation 7a8a99c8-b520-4f7e-bd11-4cfa72266b36 · outbound

This paper cites MoELoRA: Contrastive Learning Guided Mixture of Experts on Parameter-Efficient Fine-Tuning for Large Language Models.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging MoELoRA: Contrastive Learning Guided Mixture of Experts on Parameter-Efficient Fine-Tuning for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:34.475689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:34.475689Z digest=sha256:063a4131d1526708f2d6a9aa6b9f3d6453792dbdb4a3038f4af12e0be6e41b05

Observation 965601e2-00c2-4dbe-b1b5-8a9b2d8b731d · outbound

This paper cites Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:34.591129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:34.591129Z digest=sha256:c557d9846ff254fc7044e6d8293c4a078b440659f4ebb9139da966919fce1825

Observation fdbba087-fced-4277-9e75-5854c8769165 · outbound

This paper cites Orthogonal adaptation for modular customization of diffusion models.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Orthogonal adaptation for modular customization of diffusion models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:34.686559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:34.686559Z digest=sha256:007c74a9fe51b2277d455aa4c7134a19e0b868ebc7d39d5a6eb62677dcb5eb42

Observation 3ed1964a-6fa9-44ec-8605-31cba3b67616 · outbound

This paper cites Ensemble learning.Ensemble machine learning: Methods and applications, pages 1–34, 2012.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Ensemble learning.Ensemble machine learning: Methods and applications, pages 1–34, 2012

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:44:38.461282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:44:34.813147Z digest=sha256:3616b6f219fd6ac57a3d52261b3cff0a7af00066c5336084887b79bd85e9622f

Observation ae21ea70-f297-4238-ad6c-fc313818264b · outbound

This paper cites Acceleration of stochastic approximation by averaging.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Acceleration of stochastic approximation by averaging

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:34.956683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:34.956683Z digest=sha256:fce5b161670b2c059c1d584da70c6f943754e7ac004c966b5f05bc768669b4e3

Observation 28bdaed8-72f2-4824-90a6-750e342590e2 · outbound

This paper cites LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:35.077153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:35.077153Z digest=sha256:6e32cccae285f6c54ab999ba3f9d5e9a758bc098f13fb48d21745f3bde0999de

Observation 0130fcf2-28b3-47cb-a2ed-04fdd78c52a3 · outbound

This paper cites Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards.Advances in Neural Information Processing Systems, 36:71095–71134, 2023.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards.Advances in Neural Information Processing Systems, 36:71095–71134, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:35.176600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:35.176600Z digest=sha256:7bfdf7221ac0276c3bc859d7f19ea9c51066d6724afc5fc15486848b4639879f

Observation 54698c76-c14f-44bb-9fa6-b99da878fa47 · outbound

This paper cites Language Models are Multilingual Chain-of-Thought Reasoners.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Language Models are Multilingual Chain-of-Thought Reasoners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:35.245526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:35.245526Z digest=sha256:65872cec47b6711aae1cac4d0a7ba3441449afaef6439fc9e8f5f31d80211d65

Observation b91491c9-ebee-4694-bcdc-98a6d2ef8eea · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:35.393197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:35.393197Z digest=sha256:8ea8a1774b479fe05d46fdc0ae7902b2f7f696fcac6d8d9aac4d36342bcee0da

Observation edba27e7-7af2-41c2-8b25-c472e8ae6012 · outbound

This paper cites Unlocking the potential of model merging for low-resource languages.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Unlocking the potential of model merging for low-resource languages

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:44:38.282317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:44:35.501545Z digest=sha256:3e5887d9fbd7165a116e12362db8a032e8d4f6f144ae3aed24d84142a4882460

Observation 2ee1fe58-fe4d-47ae-994f-d2886fcc51d7 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Gemma 2: Improving Open Language Models at a Practical Size

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:35.591163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:35.591163Z digest=sha256:20ade4f4b5a95adea526c266af4be3fce43a3949a6c1940ae113600524472bfd

Observation a78f64a5-2263-436b-97be-29d03aaa3b5a · outbound

This paper cites Estimation in high dimensions: a geometric perspective.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Estimation in high dimensions: a geometric perspective

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:44:38.115696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:44:35.688556Z digest=sha256:02ff276f95ac88291cec2e21c8ddb709e86cbe4273f466c8bdd3b56663dac08e

Observation 6038e34c-c053-442d-8403-119bcdca394e · outbound

This paper cites Principal component analysis.Chemometrics and intelligent laboratory systems, 2(1-3):37–52, 1987.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Principal component analysis.Chemometrics and intelligent laboratory systems, 2(1-3):37–52, 1987

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:44:37.954937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:44:35.866143Z digest=sha256:0f9136309d8de65a0899fa8634f707d181f70852bdeb182131b12ad675aa3e4e

Observation 713dba1e-8030-4202-a7d1-c82ead5ede4c · outbound

This paper cites Ties-merging: Resolving interference when merging models.Advances in Neural Information Processing Systems, 36:7093–7115, 2023.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Ties-merging: Resolving interference when merging models.Advances in Neural Information Processing Systems, 36:7093–7115, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:35.949478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:35.949478Z digest=sha256:916eb861fe23869a6e44b444513778f9dd58be066371e2af0108d2e048a26824

Observation d75f9e00-ade1-4bfa-87a9-8e15c2c7303c · outbound

This paper cites What Matters for Model Merging at Scale?.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging What Matters for Model Merging at Scale?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:36.048280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:36.048280Z digest=sha256:77b0ff023ab7541faaeedbf018d870ddafd6b5200fb6db2fcb90529d024a4b7b

Observation b7b36f27-715f-4f4d-8a43-5462faa7c56e · outbound

This paper cites AdaMerging: Adaptive Model Merging for Multi-Task Learning.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging AdaMerging: Adaptive Model Merging for Multi-Task Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:36.149510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:36.149510Z digest=sha256:14f9ae543ac87d10c2fbc2e8f9b7a8bceaa75ce1803f4128a7496d91b0cb303e

Observation e76b9d63-4e76-4c68-8ac7-2615fb08f8dd · outbound

This paper cites Language models are super mario: Absorbing abilities from homologous models as a free lunch.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Language models are super mario: Absorbing abilities from homologous models as a free lunch

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:36.280683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:36.280683Z digest=sha256:e01fce8582673db7c0f9b422fd0889fd290faa72ab6fed37cb2690eec61f944d

Observation 40992453-269c-43b9-915f-861568f81fa5 · outbound

This paper cites Emotion detection on tv show transcripts with sequence- based convolutional neural networks.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Emotion detection on tv show transcripts with sequence- based convolutional neural networks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:44:37.766416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:44:36.447501Z digest=sha256:51469d753a6d39fe5bb150c5553fb91e98b8ece4d7df9fc89bb6433b562c2e21

Observation 3ecc682e-1cf1-4fd6-bd56-cee4ab56e020 · outbound

This paper cites Composing parameter-efficient modules with arithmetic operation.Advances in Neural Information Processing Systems, 36:12589–12610, 2023.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Composing parameter-efficient modules with arithmetic operation.Advances in Neural Information Processing Systems, 36:12589–12610, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:36.543531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:36.543531Z digest=sha256:c5686b73c9da5bfd0be26debc1fe30cf15ddcadfb9ea862a081cc95008df2b4f

Observation e151f6e4-f5a9-48ce-8208-125b4414a467 · outbound

This paper cites Nature-Inspired Population-Based Evolution of Large Language Models.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Nature-Inspired Population-Based Evolution of Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:36.640223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:36.640223Z digest=sha256:6ba6281436a315e0a514c0b9afc84c0ee62a3b3b197cfe14d5aca353d81f2785

Observation ae8bc4e2-e34d-47ee-aacc-744cc08cdf91 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:36.776858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:36.776858Z digest=sha256:e353a013d9b2e7a63aba193a19ce8726a1a7fa3df5d7d27490ecc65923f18abb

Observation b5c8c531-4c32-486b-8dce-fc8bdfb86d45 · outbound

This paper cites particle.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging particle

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:44:37.554033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:44:36.927598Z digest=sha256:efbf4a83e6fea7480dee2c89561c7866ea528e7c99cb60a1ff5d9596cc126fdd

Pith citing papers

Observation 6d716525-a76d-431b-822c-a9ccb32d56a8 · inbound

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities cites this paper.

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities Why Do More Experts Fail? A Theoretical Analysis of Model Merging

Reference 245

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:16:04.730252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T22:16:04.386706Z digest=sha256:5c0c7124ea7225f205d678806db2d2d139e4695a9d4f3658fb5573f8bd10f642

Observation 26ade7fe-5909-4180-85f5-c4b39f0e4243 · inbound

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging cites this paper.

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging Why Do More Experts Fail? A Theoretical Analysis of Model Merging

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:42:58.884086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T20:24:30.679219Z digest=sha256:459003c4d68ba196e4058cc3a433ade73c238d22a24c5f3312843b4d14fe9828

Observation ad5d758d-41b7-48d4-a036-7e2ace2d781f · inbound

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging cites this paper.

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging Why Do More Experts Fail? A Theoretical Analysis of Model Merging

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:05.733226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T09:12:22.712240Z digest=sha256:192db09f50831aa3fb00bf6fbd29e82741ee3cdbffbac8664d72b74de329c35f

Observation 2d3685b4-69b6-40b0-bf5d-a0e5f3951412 · inbound

Model Merging to Evolution: Parameter Space Exploration for Expert Models cites this paper.

Model Merging to Evolution: Parameter Space Exploration for Expert Models Why Do More Experts Fail? A Theoretical Analysis of Model Merging

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:24:38.146287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T11:20:04.579617Z digest=sha256:fa81db4da8316d91080412c6ec597425b8ece43ec52923c020d4cccdd268b2d3

Observation dc62f227-3112-43cc-a5f8-3d2b3c7ff4d5 · inbound

Sharpness-aware Model Merging with Salience Recovery for LLM-based Cross-Domain Sequential Recommendation cites this paper.

Sharpness-aware Model Merging with Salience Recovery for LLM-based Cross-Domain Sequential Recommendation Why Do More Experts Fail? A Theoretical Analysis of Model Merging

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T02:42:45.039852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:42:45.039852Z digest=sha256:54b8b550a10439c77e6074b53596dab3ef011eed06bb55e96348fafe23ca3708