Pith. sign in

Paper Citation Record · LEDGER

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models

As of 11 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2501.16650.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.16650 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T11:47:24.980984Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-09T21:48:48.992712Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T14:26:03.981445Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 05dc27ee-b118-424f-a936-014a0281c8c8 · outbound

This paper cites Representation Topology Divergence: A Method for Comparing Neural Network Representations.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Representation Topology Divergence: A Method for Comparing Neural Network Representations

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.789560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.789560Z digest=sha256:a9c3caf5ab0510a9c417f8ff97987f0ba899e87ac04c5abe199656071c68d9df

Observation 16b0b76c-3dd6-466b-90f2-f884eb98425a · outbound

This paper cites Define: • X ∈ Rn×m as X = [e1, e2,.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Define: • X ∈ Rn×m as X = [e1, e2,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:47:25.815710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.957502Z digest=sha256:3702f31f32885c8a76f747c502e47a5ec93a51b495bc17c9642f0517c3277e3b

Observation af0f8470-2e8d-4d00-8085-ea81dd4b6b9f · outbound

This paper cites The Llama 3 Herd of Models.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.809527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.809527Z digest=sha256:791efd4df14fc7922e2231cbf182f20ba5839774518eb7f2ffc49abf75a715a5

Observation d5fb136d-9d71-40ac-bf3a-c11c482c1e98 · outbound

This paper cites Training Compute-Optimal Large Language Models.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Training Compute-Optimal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.819179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.819179Z digest=sha256:4eee1494c7559dd7f8c72dd29332a42276b27584d3cfc1900c2f1bf64ea7ee46

Observation 0b8f7962-3c3a-42c1-bd10-c6a89c6361f1 · outbound

This paper cites an unresolved cited work.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:47:25.786022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.966835Z digest=sha256:0c934d44b1c2145dfa0df85ff19548dc526fd0df49301a1b122eb77e6a3f3510

Observation 6f03a463-269e-4898-be0b-47fa11a3778d · outbound

This paper cites Mixtral of Experts.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Mixtral of Experts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.828791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.828791Z digest=sha256:665b1203cb87045042c29c5d62a11b93f39ae75cfb20bc4d727423c9965c6d87

Observation fe4e2a0f-733d-4ae0-b424-d832c1395dae · outbound

This paper cites an unresolved cited work.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:47:25.830335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.953178Z digest=sha256:e30d4ed7b3f93f642b38a9c81a4f42abe5bec782a126128671f1f496962b46e2

Observation 2bd6e2b6-f581-4244-bbbd-fa3d2513e788 · outbound

This paper cites Similarity of Neural Network Models: A Survey of Functional and Representational Measures.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Similarity of Neural Network Models: A Survey of Functional and Representational Measures

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.843051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.843051Z digest=sha256:cf73fd24104d4d6b62406fb14645cd9b4d4c6a22301f420d7cd094549639f8f9

Observation dd5e5149-a33e-4bda-8c96-0951ce195283 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.852704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.852704Z digest=sha256:42d65fbf7a84d9e7903bf360ff5276aeeaf42c947480dbb7d22368de28a8194c

Observation 87eff07f-a74a-4e6e-afa0-0866962e880c · outbound

This paper cites Beyond KV Caching: Shared Attention for Efficient LLMs.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Beyond KV Caching: Shared Attention for Efficient LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.862817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.862817Z digest=sha256:68ddc8175f934f2c45c8997a050cc7f2f44829e3cdb772a8fcb9697560880ddb

Observation de712690-3514-4dd2-aa11-0747dd2bdc1b · outbound

This paper cites Training language models to follow instructions with human feedback.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Training language models to follow instructions with human feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.871580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.871580Z digest=sha256:777e4a125487062cf5232fc3693fd9f80941472f54124153a5fa206d024682ea

Observation 49953636-8b04-4da7-a906-40a9c50a41a3 · outbound

This paper cites SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.885060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.885060Z digest=sha256:71ea3d143c1fad54d64e50b6859bce0af6173e2627ddd05910162c2cb7d846d3

Observation 742f666a-d0fa-43df-bfe7-96b0b3c99ad2 · outbound

This paper cites Similarity of Neural Networks with Gradients.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Similarity of Neural Networks with Gradients

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-10T11:47:25.139011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.889596Z digest=sha256:b1696f9a7ee19d020595aa063dbfc1036a18397bdb5fa3210018ab51e259fe43

Observation d43e7ff9-9acc-41a8-9deb-dc256d17c6b5 · outbound

This paper cites JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.894568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.894568Z digest=sha256:796dfd41cdf3f04c13872bfde7874cbef5bfec6a869fa0fd90047b5169d4ffd3

Observation e6288e51-2e02-473e-b75b-86b6f3a86d00 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.899313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.899313Z digest=sha256:82abe2d7e250123d39004f8888e3c738c449a2de1502c230d550afce2d1f063f

Observation 6fba8e80-86cd-4fbb-8070-64940cbb7e2c · outbound

This paper cites Similarity Analysis of Contextual Word Representation Models.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Similarity Analysis of Contextual Word Representation Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-10T11:47:25.086616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.903806Z digest=sha256:840a64a55b2eeb599db16e219e576e03339e435643b2a96a97f026598aeedc03

Observation 59f177f1-2a59-4317-9d92-c4ffb39e0fe2 · outbound

This paper cites GLM-130B: An Open Bilingual Pre-trained Model.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models GLM-130B: An Open Bilingual Pre-trained Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.908781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.908781Z digest=sha256:7bf30b0c3daa89c812242680d7f4366c9bddfa5874b9b29107b18ce85727232c

Observation f0d1d419-7a09-40d2-a230-f7be7e0e84a0 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models OPT: Open Pre-trained Transformer Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.913503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.913503Z digest=sha256:738d4e0c3bd74e60fba756e3da53e755e9f7d9de99704f4543d8d6001c3c5e3b

Observation 4e66e345-a351-4244-8588-af7abb3a6c7e · outbound

This paper cites CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.918198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.918198Z digest=sha256:5799e11891fe6f8417808f22615314777d54ed1df35c7feed5f7f8dd768fd427

Observation 82837663-8add-443f-a6c0-3baf4e2fdac8 · outbound

This paper cites Taming Sparsely Activated Transformer with Stochastic Experts.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Taming Sparsely Activated Transformer with Stochastic Experts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.923354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.923354Z digest=sha256:d8510a5db1826305697aa3a8c49afe933bfe7052d418b757ee2ec28274067dc6

Observation 4a9ab880-d151-46f8-b573-141b76c9f3ac · outbound

This paper cites Let X, Y∈ Rn×m and let PX , PY ∈ Rm×m be permutation matrices.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Let X, Y∈ Rn×m and let PX , PY ∈ Rm×m be permutation matrices

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:47:25.902803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.929285Z digest=sha256:ea139f60d5067891fc1604112ae3c01a6082943f9bd7f2c236484b6b610e970c

Observation 909b7bc1-4c15-4b30-a5b1-f77afa32dbe9 · outbound

This paper cites −0.6676 0 .5171 −0.5357 −0.7310 −0.5917 0 .3399 −0.1412 0 .6185 0 .7730 # , Y =.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models −0.6676 0 .5171 −0.5357 −0.7310 −0.5917 0 .3399 −0.1412 0 .6185 0 .7730 # , Y =

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:47:25.888325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.934700Z digest=sha256:61ab0f02bc6528a98a59f612fccefdb052d5db3ae7ca5034174548693a1642c0

Observation 0cd29245-4d5f-40a1-a919-0e4fa3d0ccf7 · outbound

This paper cites an unresolved cited work.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T11:47:25.873425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.939478Z digest=sha256:804f2190e37d5f26b2f8e73434f51f0188377d6d1c2dd29b47c4949c0b9516f4

Observation 82a0312c-b81d-41ba-b9e4-2206007b2ecc · outbound

This paper cites Meanwhile, TX and TY denote truncated identity matrices that retain the left singular vectors, ensuring that the accumulated variance meets a predefined limit.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Meanwhile, TX and TY denote truncated identity matrices that retain the left singular vectors, ensuring that the accumulated variance meets a predefined limit

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:47:25.859383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.943871Z digest=sha256:65946ce2c051f815795b942208584836e010311fe364203970edc148bd849b95

Observation 44cb9398-508f-492b-bc0d-891e6435ab4d · outbound

This paper cites Meanwhile, TX and TY denote truncated identity matrices that retain the left singular vectors, ensuring that the accumulated variance meets a predefined limit.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Meanwhile, TX and TY denote truncated identity matrices that retain the left singular vectors, ensuring that the accumulated variance meets a predefined limit

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:47:25.844812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.948575Z digest=sha256:5be2082df94d1f0e8f30e33aad1290fa9e5b6eb8ba66451fce1d72b65f8bd36d

Observation b61dc0a0-6574-43be-b91c-8f9d3dbf5ba2 · outbound

This paper cites ∥X − Y ∥F = √ 2m = Ω(√n).

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models ∥X − Y ∥F = √ 2m = Ω(√n)

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:47:25.800069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.962147Z digest=sha256:0b1a406199ed066c285f2a2acc96850aa687caba96400d2f49f0205341c15c00

Observation 988ec9d1-67a7-463f-999f-558e392f1c38 · outbound

This paper cites This formulation indicates that a larger Off-Diagonal Average Cosine Similarity value corresponds to a lower degree of orthogonality in the matrix.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models This formulation indicates that a larger Off-Diagonal Average Cosine Similarity value corresponds to a lower degree of orthogonality in the matrix

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:47:25.771907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.971627Z digest=sha256:5fd08f26aac11e444f0713173b94fb08d9dee6575587a02536fba11f23bcaebd

Observation 32006ad6-cb52-4cc2-8ffb-0f74e8a5f83b · outbound

This paper cites To quantify this, we define the similarity ratio as the ratio of the similarity scores between models (A) and (B) to those between models (A) and (C).

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models To quantify this, we define the similarity ratio as the ratio of the similarity scores between models (A) and (B) to those between models (A) and (C)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:47:25.757167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.976246Z digest=sha256:8a7a5036a6fdc9d52370c38b692c8c2200d14825ed3caa2e313f145bcb4c0808

Observation 633553c8-72a7-4ec2-bef5-ccbf2e0222bd · outbound

This paper cites Importance of the Maximization Function The maximization operation in the M AXCOSSIM function plays a crucial role in the DOCS algo- rithm.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Importance of the Maximization Function The maximization operation in the M AXCOSSIM function plays a crucial role in the DOCS algo- rithm

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:47:25.742747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.980984Z digest=sha256:1d408fcfd323957865e61e2152d1e4c489dec3170b56cfca4114c1601f1699e4

Observation 386761bf-9eb5-4fbd-a5e8-921345391936 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 1984

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.880626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.880626Z digest=sha256:a4f4c925e84efd046c3e3d6246a6a7b1814fce44253d18745a92115cf63e665f

Observation fbfb80ff-3275-4797-99fa-f4c9b4098f7c · outbound

This paper cites Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.814636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.814636Z digest=sha256:6bda840a680de7131e3fa12c95a29972e37f7a0fdf6c1e712799b409c28c80bd

Observation 6d714a17-3487-4d22-9899-f7490ef1e472 · outbound

This paper cites The Remarkable Robustness of LLMs: Stages of Inference?.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models The Remarkable Robustness of LLMs: Stages of Inference?

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.848101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.848101Z digest=sha256:b8fc22b4017b2e74f039e89da5680cd6434fee58bd8c2edaaa7ea4d808016741

Observation d149ff03-2ff1-4033-bdc5-b3c1de397a85 · outbound

This paper cites StableMoE: Stable Routing Strategy for Mixture of Experts.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models StableMoE: Stable Routing Strategy for Mixture of Experts

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.804883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.804883Z digest=sha256:5fbc5bbc8f108d4445c86c23189923f8d1d0700967ef3a65d56db6a2848ed24f

Observation 9e1dd2c9-96f1-4a1e-a00d-833fa9c4a5a6 · outbound

This paper cites Contrasim–analyzing neural representations based on con- trastive learning.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Contrasim–analyzing neural representations based on con- trastive learning

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:47:25.917187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.876404Z digest=sha256:8b92a0029a8db98b5e00ec54ef63baa77fd338824a4bebb3dc0e828539938e18

Observation b9d9a71b-ec28-4fee-b29c-c55482b03551 · outbound

This paper cites Cross-layer attention sharing for large language models.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Cross-layer attention sharing for large language models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.867353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.867353Z digest=sha256:8630b0d62f52503ce581fd59eebb14b6be32f9f1b62250d126299c243650660e

Observation 1dd28962-02a1-440f-a866-482422719b4d · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.824244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.824244Z digest=sha256:83ff55ad5a0874d24f6db363c784ba5ff7f5a1a93532a8222c2870e4ef5fff64

Observation b3c8182b-3ff7-4c20-929b-79fa92b1f6ad · outbound

This paper cites FLM-101B: An Open LLM and How to Train It with $100K Budget.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models FLM-101B: An Open LLM and How to Train It with $100K Budget

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.857799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.857799Z digest=sha256:3ddfb4be0ef532c1b20330dc405a913f4b4652dc6065c50d87d3499ab2c6f841

Observation 0c82ae55-6a89-4f9e-a0d1-7af428c33d55 · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.795182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.795182Z digest=sha256:bdd6d8a9a871a80366a456355e302e112959f3a9728e3d96e9d1601c7b63be6b

Observation 434f60a1-e965-4f14-9b1d-58d4a473bd9f · outbound

This paper cites Language models are few-shot learners.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Language models are few-shot learners

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T11:47:25.932167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.799923Z digest=sha256:15de9aac58d2dfb1a76a0df4a2075e30c5b063b4050b2810739b9bb265a1edcf

Observation d97dc3fe-4a23-4d6c-b2eb-37887b9e14ce · outbound

This paper cites Scaling Laws for Neural Language Models.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models Scaling Laws for Neural Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T11:47:24.837899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:47:24.837899Z digest=sha256:77e01db8447364fda9097704c90d3671fb1c3aeafc04fcb5cdda31faaa2b46ea

Observation 8f909818-9c3f-43f9-8c4b-ef2b3b73bbe3 · outbound

This paper cites NeCo@ALQAC 2023: Legal Domain Knowledge Acquisition for Low-Resource Languages through Data Enrichment.

DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models NeCo@ALQAC 2023: Legal Domain Knowledge Acquisition for Low-Resource Languages through Data Enrichment

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T11:47:25.612703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T11:47:24.833304Z digest=sha256:949b445d69e95b00a2d730bc994ea54b8eb6354daaf7d8c22365c08e6a60092e

Pith citing papers

Observation d68b3489-039f-49cc-b201-54ad94a0c68d · inbound

Low-Rank Adaptation Redux for Large Models cites this paper.

Low-Rank Adaptation Redux for Large Models DOCS: Quantifying Weight Similarity for Deeper Insights into Large Language Models

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:26:03.983869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T21:48:48.992712Z digest=sha256:a2467733e4b287dff1d78db564f78ffdd05f941e6e0775342fa8d3d35d4cf6de