Pith. sign in

Paper Citation Record · LEDGER

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization

As of 5 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2605.08809.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.08809 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:33:45.192139Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact13
  • verified fuzzy3
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0e0be72-5092-499a-9868-616165665746 · outbound

This paper cites Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:28.852066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:79006376dfd9cec88180c981ac5e2e53ba251f334c11bf3048a5b477e462d779

Observation 54955a1a-11ae-4f2a-9434-07f95f53ab40 · outbound

This paper cites Joint selection for large-scale pre-training data via policy gradient-based mask learning.arXiv preprint arXiv:2512.24265.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization Joint selection for large-scale pre-training data via policy gradient-based mask learning.arXiv preprint arXiv:2512.24265

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:28.903351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:37cbb7c79a3b7c4095833029093da3db20df4a16fb58c03be16560a67f82a622

Observation 1db87a18-273a-4772-87e1-aa024fc05650 · outbound

This paper cites He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization He, B., Yin, L., Zhen, H.-L., Liu, S., Wu, H., Zhang, X., Yuan, M., and Ma, C

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:28.917670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:57e9d89cb706eb8230511e471c7c1bc5b58cd6790d906f33e780b8868beb6475

Observation 97faa1aa-aee3-4e3f-a323-d6b70a290b16 · outbound

This paper cites SimCSE: Simple Contrastive Learning of Sentence Embeddings.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization SimCSE: Simple Contrastive Learning of Sentence Embeddings

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:48:23.738089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:f1935a506fa2f555fa2cc7f5c952d6edacd46bc918b9964a9351870b8acc6f4f

Observation 8a831358-8906-4b57-a731-761465cacc49 · outbound

This paper cites Jun Hu, Wenwen Xia, Xiaolu Zhang, Chilin Fu, Weichang Wu, Zhaoxin Huan, Ang Li, Zuoli Tang, and Jun Zhou.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization Jun Hu, Wenwen Xia, Xiaolu Zhang, Chilin Fu, Weichang Wu, Zhaoxin Huan, Ang Li, Zuoli Tang, and Jun Zhou

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:11:48.180033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:9c53a7efce6f374d5bc2463fd11160571974df69cceef05c9e849ca2a19b396b

Observation 8e02469a-04b3-4fc2-ac2e-a3023d04441a · outbound

This paper cites Mixtral of Experts.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization Mixtral of Experts

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:16:28.913257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:97a9c4ed701108730b54b2dfaf703853da7844c76b1fd4d0adecbba151968d83

Observation f2855f91-1f3f-48db-b801-98d481d39e39 · outbound

This paper cites A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:28.897313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:dbdcf04d4b2996d391aa885e38e56a549bde0d89bed617ba42f718180b187eb9

Observation d09df841-1e59-46a7-9b1c-7307c3208f27 · outbound

This paper cites Text and Code Embeddings by Contrastive Pre-Training.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization Text and Code Embeddings by Contrastive Pre-Training

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:24:12.120554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:25225aec9ebbca01533f59266953175cef7ccd87348f330902d20a1848950352

Observation ffe8bed3-18da-4e29-8c19-9f7169094d3a · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization Representation Learning with Contrastive Predictive Coding

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:16:28.820669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:6ddf12f553ab943dca0756de47b029e2b27f440d3fb6d3592d50d44d4b1fb333

Observation 61134f1e-7314-4cdf-ad7e-aa4ece46efe2 · outbound

This paper cites GLU Variants Improve Transformer.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization GLU Variants Improve Transformer

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:16:28.892187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:ea87558d0581759c9a06e14c5219fcd0205f72dc6771e1c2fc778e215253c535

Observation 9e249d79-f9f5-475d-98a4-ed5441191f7c · outbound

This paper cites A simple contrastive learning framework for interactive argument pair identification via argument-context extraction.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization A simple contrastive learning framework for interactive argument pair identification via argument-context extraction

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:11:48.184269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:71c85362daf114d28ede7a7cb1d11eb428e066b5e67fd7cf83199310ac47c794

Observation 1018947d-9799-41ad-96ba-0ab476a704d6 · outbound

This paper cites MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:54:43.618312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:66114e0389350f210434b0de9e846c7f4d075686ccccd9c93baab5e41b6e1a08

Observation f2e5d0f5-5471-49db-8328-91c0c4e4ecad · outbound

This paper cites LLMs are Also Effective Embedding Models: An In-depth Overview.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization LLMs are Also Effective Embedding Models: An In-depth Overview

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:28.880514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:fb39964fd3fb8dc4d5195d5e4a241ae8df8be169fd8668d5748fc227961643dc

Observation 85cc35bc-1a59-4504-b3bd-f1b842046675 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:16:28.863749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:32eec6771c47aee823f025920c85126bea7e0a3a8105ad765433c05d169e4aa0

Observation 808990e5-6e2e-493a-889d-8d20a7c105c5 · outbound

This paper cites AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-10T02:11:12.038297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:c9c9e700d7688b0dea2f5adacb70175c0bb2ca0b1d7a483b4969ae3b763f4a04

Observation e1f8b234-50ae-4a35-a057-d5b760ef3c42 · outbound

This paper cites Do Generated Data Always Help Contrastive Learning?.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization Do Generated Data Always Help Contrastive Learning?

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:28.838582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:fc2dea1fafbe5488e8ca73214256bc0ccae9d313d6f53227d23351f0eeb4409d

Observation 688420ce-fd93-446a-ae48-f658e60628ec · outbound

This paper cites Is contrastive learning necessary? a study of data augmentation vs contrastive learning in sequential recommen- dation.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization Is contrastive learning necessary? a study of data augmentation vs contrastive learning in sequential recommen- dation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T19:11:48.188493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:827fdd80c10d4e9187642e86b4288114320162ab190575ab96b4cd8f43afda5b

Observation e7fab80f-0bf7-4832-9f40-e7fadc40ed2c · outbound

This paper cites an unresolved cited work.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-12T19:11:48.211340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:a05442d03e21d117e2fab5eaff969a0c5f8f1e0f82225b12a7098cd23232025e

Observation 740bbdbc-b5ea-4b88-ba79-11c904593bca · outbound

This paper cites an unresolved cited work.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-12T19:11:48.192602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:bd1148c4cc0324992c4baa2dc13df04a806b390d819ef8505ce4a37f83868b11

Observation 3e4bed0a-3173-4bd3-8da4-40e5dfdcdd75 · outbound

This paper cites an unresolved cited work.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-12T19:11:48.196449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:453c51db65f005d196b0451d7f7dbb8412b4ea107c219291c3f24b8619434827

Observation e97ad2c6-7ff8-4069-af58-49c27bc94206 · outbound

This paper cites an unresolved cited work.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization Unresolved cited work

Reference 21

Resolution
malformed identifier
raw_fallback, observed 2026-05-12T19:11:48.207258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:6bbc6270cd21e180bfba6dafd6e92d2b4a8aa4fac0040ae93f1155aa001e2e47

Pith citing papers

No inbound Pith citation observations are available.