Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:46:28.238597Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 11 inbound Pith citation observations for arXiv:2505.12082.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:46:28.238597Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:30:10.163726Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:49:46.245243Z
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dc5cf18c-451a-4f70-9c22-712ccb546981 · outbound
Model Merging in Pre-training of Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 113a2c22-078a-44da-9bb5-3ac6372e80b4 · outbound
Model Merging in Pre-training of Large Language Models Evolutionary optimization of model merging recipes
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06fe965-8853-4ab4-9c45-54f33b7cc8c9 · outbound
Model Merging in Pre-training of Large Language Models Program Synthesis with Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bcbde28-ffdc-4097-9834-52c36772cf39 · outbound
Model Merging in Pre-training of Large Language Models DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abf61e49-c49b-4a58-970b-e6ea3f5c7287 · outbound
Model Merging in Pre-training of Large Language Models Evaluating Large Language Models Trained on Code
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c91f5bf3-8d9e-48fc-94aa-93bffa4db82e · outbound
Model Merging in Pre-training of Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1657b69-436f-4b9a-b7ab-d23823604a78 · outbound
Model Merging in Pre-training of Large Language Models Training Verifiers to Solve Math Word Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8d5fa3a-2383-49ef-ae99-9e16c020f42b · outbound
Model Merging in Pre-training of Large Language Models Gradient descent on neural networks typically occurs at the edge of stability
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f6af6af-c17b-452e-9cb9-7cdb04977025 · outbound
Model Merging in Pre-training of Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 694e03df-c1b9-420e-9715-8601d83ca4e1 · outbound
Model Merging in Pre-training of Large Language Models DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation afdfec0b-13eb-410b-90a2-86a92e6d70e3 · outbound
Model Merging in Pre-training of Large Language Models The llama 3 herd of models.arXiv e-prints, pages arXiv–2407, 2024
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d02f92f1-ae8a-4335-b87b-c16443f94f05 · outbound
Model Merging in Pre-training of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fdbbd56-4a27-4c7a-a7f2-d1535cbc2ff4 · outbound
Model Merging in Pre-training of Large Language Models Scaling laws and compute-optimal training beyond fixed training durations.Advances in Neural Information Processing Systems, 37:76232–76264, 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 161b767d-48e6-4e61-939a-a37620cbbca0 · outbound
Model Merging in Pre-training of Large Language Models Measuring massive multitask language understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7168105-42fe-4f7b-be6e-5506274842cf · outbound
Model Merging in Pre-training of Large Language Models MiniCPM: Unveiling the potential of small language models with scalable training strategies
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4370adf0-5a6c-4cb9-937f-865e0524054d · outbound
Model Merging in Pre-training of Large Language Models C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d96def94-1097-4ab3-90da-66a2037988e4 · outbound
Model Merging in Pre-training of Large Language Models The exponentially weighted moving average.Journal of quality technology, 18(4):203–210, 1986
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caade15a-3dbd-4e92-934c-7eeadf223440 · outbound
Model Merging in Pre-training of Large Language Models Editing models with task arithmetic
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 15e89984-c6ad-4484-935f-64099d86899e · outbound
Model Merging in Pre-training of Large Language Models Dataless knowledge fusion by merging weights of language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2429eb69-8451-44db-801b-b86f1afe2439 · outbound
Model Merging in Pre-training of Large Language Models Some properties of a simple moving average when applied to forecasting a time series.Journal of the Operational Research Society, 50(12):1267–1271, 1999
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a7fe01eb-919e-48e0-a51d-2a8c413791fa · outbound
Model Merging in Pre-training of Large Language Models PopulAtion Parameter Averaging (PAPA)
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac51e743-4037-4e1c-b1bf-25e121691780 · outbound
Model Merging in Pre-training of Large Language Models TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2b0de40c-ddfa-430a-8e69-f771b2f930d0 · outbound
Model Merging in Pre-training of Large Language Models Stop wasting my time! saving days of imagenet and BERT training with latest weight averaging
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6e4a93c5-356a-4446-9a3a-9e46f31c46e4 · outbound
Model Merging in Pre-training of Large Language Models Scaling Laws for Neural Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e54e9882-bdcf-4eaa-9728-ef474e5c1122 · outbound
Model Merging in Pre-training of Large Language Models Trainable weight averaging: Efficient training by optimizing historical solutions
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 08235f68-b1ff-48d8-961b-a118c409527e · outbound
Model Merging in Pre-training of Large Language Models DeepSeek-V3 Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceb07be9-756a-4a49-a66e-d7565e759fcd · outbound
Model Merging in Pre-training of Large Language Models Checkpoint Merging via Bayesian Optimization in LLM Pretraining
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3309126d-ca29-4e54-a4e6-7c4fee9cfeb0 · outbound
Model Merging in Pre-training of Large Language Models SGDR: Stochastic gradient descent with warm restarts
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcf1b735-1106-4bb8-a3d5-e7a6358126f7 · outbound
Model Merging in Pre-training of Large Language Models Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 13cb1706-2de0-43c7-a4af-c32ebba4822d · outbound
Model Merging in Pre-training of Large Language Models Merging models with fisher-weighted averaging.Advances in Neural Information Processing Systems, 35:17703–17716, 2022
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcb3d693-7a59-41d3-937b-9a222a560c6c · outbound
Model Merging in Pre-training of Large Language Models An Empirical Model of Large-Batch Training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0cfc1d1-ea07-4974-a2ab-daedd6058278 · outbound
Model Merging in Pre-training of Large Language Models The weighted moving average technique.Wiley Encyclopedia of Operations Research and Management Science, 2010
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 15e7104d-207f-4cdc-af77-ce5ef72d3731 · outbound
Model Merging in Pre-training of Large Language Models Gpqa: A graduate-level google-proof q&a benchmark
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6017fe27-3ec2-46c4-acce-6ac0f84d3535 · outbound
Model Merging in Pre-training of Large Language Models Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f733f241-f864-4aa3-b47b-e00bc7be65b6 · outbound
Model Merging in Pre-training of Large Language Models Early weight averaging meets high learning rates for LLM pre-training
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ab4f75e6-a192-43b5-982f-1f1bf5f31dd0 · outbound
Model Merging in Pre-training of Large Language Models Seed-thinking-v1
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e384067-98dc-4227-a9bd-81e473be7b7e · outbound
Model Merging in Pre-training of Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a370094-ffcc-4ee8-ae56-bd33208f8ad3 · outbound
Model Merging in Pre-training of Large Language Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1205ee6d-ee6c-4766-bad1-f09d51b5f3ee · outbound
Model Merging in Pre-training of Large Language Models Challenging BIG-bench tasks and whether chain-of- thought can solve them
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 73f8b1d7-a9a4-479b-a23e-544e4c0951be · outbound
Model Merging in Pre-training of Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c6fd893-5ccb-4866-8714-d9865e9d7e1a · outbound
Model Merging in Pre-training of Large Language Models Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3f52f625-1ed8-4e1d-b7c6-b32076afabdb · outbound
Model Merging in Pre-training of Large Language Models Dai, and Quoc V Le
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8b17a5cd-4a8d-4f45-bece-510fea2ae4d2 · outbound
Model Merging in Pre-training of Large Language Models Livebench: A challenging, contamination-limited LLM benchmark
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5ed6c65a-7216-4061-91ea-c214d88deb82 · outbound
Model Merging in Pre-training of Large Language Models Small-scale proxies for large-scale transformer training instabilities
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 07f1acda-f168-41bb-8595-4c7eb99b71b7 · outbound
Model Merging in Pre-training of Large Language Models WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3f9a73ed-d049-4584-b97b-64c716ee8d7e · outbound
Model Merging in Pre-training of Large Language Models TIES-merging: Resolving interference when merging models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 975fafc3-1f35-4e9e-acd6-997afcb40bc5 · outbound
Model Merging in Pre-training of Large Language Models Baichuan 2: Open Large-scale Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1419f775-e006-4a89-8060-f18f0bf6a6ae · outbound
Model Merging in Pre-training of Large Language Models Qwen2.5 Technical Report
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20d6aab7-72a0-45d2-a645-5d0e0a4bd855 · outbound
Model Merging in Pre-training of Large Language Models Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa8c7e5b-22e5-4888-959e-d68ffeec81a2 · outbound
Model Merging in Pre-training of Large Language Models Adamerging: Adaptive model merging for multi-task learning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bc2a334e-9ee0-4af7-82c9-269437af24c5 · outbound
Model Merging in Pre-training of Large Language Models Language models are super mario: Absorbing abilities from homologous models as a free lunch
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3e4cb3b1-6dcc-4ef7-9ac8-a149d50e25fe · outbound
Model Merging in Pre-training of Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa64e4dd-c4bf-4bef-8fa8-94024c6853c2 · outbound
Model Merging in Pre-training of Large Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d7d8925-ac6d-4eca-9cb6-31086e773e10 · outbound
Model Merging in Pre-training of Large Language Models HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e386bee-32b1-4cea-863e-91acaf16e72a · outbound
Model Merging in Pre-training of Large Language Models Ape210K: A Large-Scale and Template-Rich Dataset of Math Word Problems
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6694c25b-3f3b-41a3-bd03-50540bfc15eb · outbound
Model Merging in Pre-training of Large Language Models AGIEval: A human-centric benchmark for evaluating foundation models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bf9d46d5-0779-438b-aab2-fd7d94bc92b3 · outbound
Model Merging in Pre-training of Large Language Models MetaGPT: Merging large language models using model exclusive task arithmetic
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 11f46116-417b-46e2-9a8f-da9c9ba887cd · inbound
Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities Model Merging in Pre-training of Large Language Models
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 04dde608-7a4e-4db6-9a65-e5d8367a23ee · inbound
Kwai Keye-VL Technical Report Model Merging in Pre-training of Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b01d7e69-effd-4359-869c-aaa6507bee03 · inbound
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training Model Merging in Pre-training of Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d532f56f-1d36-4e14-a307-9d752f503eed · inbound
ReaLM: Reflection-Enhanced Autonomous Reasoning with Small Language Models Model Merging in Pre-training of Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a559c1-04ee-4264-b5dd-27d1d9222cd3 · inbound
Waver: Wave Your Way to Lifelike Video Generation Model Merging in Pre-training of Large Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdfded52-de3c-4810-8fbc-99db82938247 · inbound
ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models Model Merging in Pre-training of Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e7bd2b19-81d5-452c-b607-0ede6828ea68 · inbound
Anytime Training with Schedule-Free Spectral Optimization Model Merging in Pre-training of Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6cac034e-95b0-4736-a464-d13520d85f90 · inbound
Kwai Keye-VL-2.0 Technical Report Model Merging in Pre-training of Large Language Models
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a161417f-12aa-443a-9431-55420b5ce9fc · inbound
Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI Model Merging in Pre-training of Large Language Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 556c49c3-1768-4df0-bda2-1621d33479bb · inbound
LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning Model Merging in Pre-training of Large Language Models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8a40e590-512f-447e-9ff7-8ec128222178 · inbound
Optimizing Visual Generative Models via Distribution-wise Rewards Model Merging in Pre-training of Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.