Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T11:47:24.980984Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2501.16650.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T11:47:24.980984Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-09T21:48:48.992712Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T14:26:03.981445Z
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 05dc27ee-b118-424f-a936-014a0281c8c8 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Representation Topology Divergence: A Method for Comparing Neural Network Representations
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16b0b76c-3dd6-466b-90f2-f884eb98425a · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Define: • X ∈ Rn×m as X = [e1, e2,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation af0f8470-2e8d-4d00-8085-ea81dd4b6b9f · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5fb136d-9d71-40ac-bf3a-c11c482c1e98 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Training Compute-Optimal Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b8f7962-3c3a-42c1-bd10-c6a89c6361f1 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6f03a463-269e-4898-be0b-47fa11a3778d · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Mixtral of Experts
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe4e2a0f-733d-4ae0-b424-d832c1395dae · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2bd6e2b6-f581-4244-bbbd-fa3d2513e788 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Similarity of Neural Network Models: A Survey of Functional and Representational Measures
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd5e5149-a33e-4bda-8c96-0951ce195283 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87eff07f-a74a-4e6e-afa0-0866962e880c · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Beyond KV Caching: Shared Attention for Efficient LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de712690-3514-4dd2-aa11-0747dd2bdc1b · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Training language models to follow instructions with human feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49953636-8b04-4da7-a906-40a9c50a41a3 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 742f666a-d0fa-43df-bfe7-96b0b3c99ad2 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Similarity of Neural Networks with Gradients
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d43e7ff9-9acc-41a8-9deb-dc256d17c6b5 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6288e51-2e02-473e-b75b-86b6f3a86d00 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fba8e80-86cd-4fbb-8070-64940cbb7e2c · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Similarity Analysis of Contextual Word Representation Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 59f177f1-2a59-4317-9d92-c4ffb39e0fe2 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models GLM-130B: An Open Bilingual Pre-trained Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0d1d419-7a09-40d2-a230-f7be7e0e84a0 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models OPT: Open Pre-trained Transformer Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e66e345-a351-4244-8588-af7abb3a6c7e · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82837663-8add-443f-a6c0-3baf4e2fdac8 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Taming Sparsely Activated Transformer with Stochastic Experts
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a9ab880-d151-46f8-b573-141b76c9f3ac · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Let X, Y∈ Rn×m and let PX , PY ∈ Rm×m be permutation matrices
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 909b7bc1-4c15-4b30-a5b1-f77afa32dbe9 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models −0.6676 0 .5171 −0.5357 −0.7310 −0.5917 0 .3399 −0.1412 0 .6185 0 .7730 # , Y =
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0cd29245-4d5f-40a1-a919-0e4fa3d0ccf7 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 82a0312c-b81d-41ba-b9e4-2206007b2ecc · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Meanwhile, TX and TY denote truncated identity matrices that retain the left singular vectors, ensuring that the accumulated variance meets a predefined limit
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 44cb9398-508f-492b-bc0d-891e6435ab4d · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Meanwhile, TX and TY denote truncated identity matrices that retain the left singular vectors, ensuring that the accumulated variance meets a predefined limit
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b61dc0a0-6574-43be-b91c-8f9d3dbf5ba2 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models ∥X − Y ∥F = √ 2m = Ω(√n)
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 988ec9d1-67a7-463f-999f-558e392f1c38 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models This formulation indicates that a larger Off-Diagonal Average Cosine Similarity value corresponds to a lower degree of orthogonality in the matrix
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 32006ad6-cb52-4cc2-8ffb-0f74e8a5f83b · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models To quantify this, we define the similarity ratio as the ratio of the similarity scores between models (A) and (B) to those between models (A) and (C)
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 633553c8-72a7-4ec2-bef5-ccbf2e0222bd · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Importance of the Maximization Function The maximization operation in the M AXCOSSIM function plays a crucial role in the DOCS algo- rithm
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 386761bf-9eb5-4fbd-a5e8-921345391936 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 1984
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbfb80ff-3275-4797-99fa-f4c9b4098f7c · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change
Reference 2005
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d714a17-3487-4d22-9899-f7490ef1e472 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models The Remarkable Robustness of LLMs: Stages of Inference?
Reference 2008
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d149ff03-2ff1-4033-bdc5-b3c1de397a85 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models StableMoE: Stable Routing Strategy for Mixture of Experts
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e1dd2c9-96f1-4a1e-a00d-833fa9c4a5a6 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Contrasim–analyzing neural representations based on con- trastive learning
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b9d9a71b-ec28-4fee-b29c-c55482b03551 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Cross-layer attention sharing for large language models
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dd28962-02a1-440f-a866-482422719b4d · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3c8182b-3ff7-4c20-929b-79fa92b1f6ad · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models FLM-101B: An Open LLM and How to Train It with $100K Budget
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c82ae55-6a89-4f9e-a0d1-7af428c33d55 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 434f60a1-e965-4f14-9b1d-58d4a473bd9f · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Language models are few-shot learners
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d97dc3fe-4a23-4d6c-b2eb-37887b9e14ce · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Scaling Laws for Neural Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f909818-9c3f-43f9-8c4b-ef2b3b73bbe3 · outbound
DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models NeCo@ALQAC 2023: Legal Domain Knowledge Acquisition for Low-Resource Languages through Data Enrichment
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d68b3489-039f-49cc-b201-54ad94a0c68d · inbound
Low-Rank Adaptation Redux for Large Models DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.