Pith. sign in

Paper Citation Record · LEDGER

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training

As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2506.10952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10952 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:19:11.124746Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T00:56:04.958757Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T00:58:25.710039Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact3
  • verified fuzzy21
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5150f381-6e00-4937-a600-9f7df201e13c · outbound

This paper cites Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:04.749969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:04.749969Z digest=sha256:5d3b74a2d0f4d526766c5cb665fb218ddfc90992fda52ef5c3e003803caacb54

Observation 20a927a0-ef93-482f-9dce-ace099b5b482 · outbound

This paper cites and Vassilvitskii, S.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training and Vassilvitskii, S

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:19.885756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:04.830069Z digest=sha256:adac34a5913a3db85340e61388d414d30bb5147d153f002017fe598c9d95128c

Observation e622fe09-9584-446d-8696-f8235fbc7bf7 · outbound

This paper cites L., Gao, J., and Choi, Y.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training L., Gao, J., and Choi, Y

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:19.575514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:04.986925Z digest=sha256:a568dc27552d0da12a6f7badfc2c4a3815208595c6f72d164147de61fb213381

Observation ca118c46-9f2f-4d5f-8f25-5724cf2b0f37 · outbound

This paper cites Cross-Table Pretraining towards a Universal Function Space for Heterogeneous Tabular Data.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Cross-Table Pretraining towards a Universal Function Space for Heterogeneous Tabular Data

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:19:12.832198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:05.168093Z digest=sha256:5345bdcd9addf88451a0b02e42986e8111c3c5044e0312cb70958663934a4144

Observation 6e43ce2c-2b94-47ff-88f7-1355d77d3822 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.292132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.292132Z digest=sha256:e260c6bbf4708af7b6943944fe8d586dd4c5ab127d98f2619e70cb57ef930e27

Observation b67c0f7c-b241-45bd-9131-ad2d9b7e24a1 · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.428897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.428897Z digest=sha256:3b1fa049f72523d2674c1f40d27042f102bd49144097cb09c4fac370a52614e0

Observation 4fc679f5-287e-45e9-84f6-f825520e8cd5 · outbound

This paper cites DOGE : Domain reweighting with generalization estimation.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training DOGE : Domain reweighting with generalization estimation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:19.238643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:05.546635Z digest=sha256:80965df54406f0dd084a00a99fdd7811c80f7fd1bb2d1694638b6c6e03bbef61

Observation d7b03bc9-e2ae-4f46-a8bf-54061c6bc317 · outbound

This paper cites Unearthing Large Scale Domain-Specific Knowledge from Public Corpora.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Unearthing Large Scale Domain-Specific Knowledge from Public Corpora

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.673354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.673354Z digest=sha256:75065f590dfb63e810c53f605cbe975f6a843307b8c18aa75620606591523f60

Observation 795ba27a-4d48-4890-8f46-906a2005cd65 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.821830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.821830Z digest=sha256:e4682b50e223a19d4e267282fab58bd394f505a2798edeaf3d3f7bcf26777083

Observation eb5ad669-b4d0-44a5-9ca1-045173519542 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training A framework for few-shot language model evaluation, 07 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.974554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.974554Z digest=sha256:57405b806b43eb222c2e27a4d55b84fb6fc1880ca6dec562ce31624cfd779fc6

Observation ee14808c-4cd9-46ba-ba7e-91720ffc2156 · outbound

This paper cites BiMix: A Bivariate Data Mixing Law for Language Model Pretraining.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training BiMix: A Bivariate Data Mixing Law for Language Model Pretraining

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:06.110067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:06.110067Z digest=sha256:22ab1e7ece2295fb8f5954a5dae3bfc726fb9a50f5204eb4ed643481d955b879

Observation 6b3fe130-c88d-48fb-acfe-7a3ef6219159 · outbound

This paper cites S em E val-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training S em E val-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:18.773659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:06.244720Z digest=sha256:1280e1f1b64570bc6da41331cff3893e29228daadee7ab0977b824489c2e9ff8

Observation 1a311c98-2493-46fa-a0c2-74b275e9c9fd · outbound

This paper cites The Llama 3 Herd of Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:06.370050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:06.370050Z digest=sha256:f1fde813692a36ed6df9f9d843738b672a56d3a67eb5945d5f520d768922a917

Observation 4fa1f33a-c468-46fb-8b96-451521052e48 · outbound

This paper cites CMR scaling law: Predicting critical mixture ratios for continual pre-training of language models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training CMR scaling law: Predicting critical mixture ratios for continual pre-training of language models

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T04:19:11.583751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:06.514933Z digest=sha256:5207e0527675fb0bc7e515ca9d7c9aa38a1e6b1d951bd7f0f3bbed8e03d49447

Observation c9b8c866-aae0-4ab3-bb09-4e4a55895a47 · outbound

This paper cites Data Selection via Optimal Control for Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Data Selection via Optimal Control for Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:06.688345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:06.688345Z digest=sha256:d3946fe4586b38cf787bd717b250e2d9c947a13b452dc73fce161c9deb1f04d9

Observation 3037b8d0-57e6-4420-bf01-73a4c64c7fc2 · outbound

This paper cites V., and Smith, K.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training V., and Smith, K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:18.402051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:06.781962Z digest=sha256:bc9f5b6d2f4818ed22dcf0e0d784e46bb1ecfda8d8f21630c9a01a6fdc50a74d

Observation bc0ecc10-2344-4fb4-bd57-c67cda0083c9 · outbound

This paper cites H., and Friedman, J.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training H., and Friedman, J

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:06.947572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:06.947572Z digest=sha256:1df3f54f1a08134ebbf5268a5ecba8a70154eaac6e313de3fca4b966a5be9cd0

Observation 2b4c0efb-c935-4e9f-9c93-bdc89412ff13 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Training Compute-Optimal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.085550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.085550Z digest=sha256:19d24db81aeca01a8b76e9d65c9afd12336389fb90d2d5f820e58c78e8e0888e

Observation a61d6b29-1bab-42bc-8290-daa94f7ee6d0 · outbound

This paper cites A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:18.067075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:07.146299Z digest=sha256:5af2adb841e281d63f77d8b1812439034b741032caadfd6f47c1d3e76fab5981

Observation b0aca12f-c72f-4a13-a262-ec07c2960b10 · outbound

This paper cites an unresolved cited work.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.273463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.273463Z digest=sha256:2bf2d33e0157573de421ec172e441dc2216a77cc1194268eabfcc35021dfdf30

Observation 5db010b9-0585-4256-a111-bd59bea0bf09 · outbound

This paper cites S., Schmidt-Thieme, L., and Grabocka, J.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training S., Schmidt-Thieme, L., and Grabocka, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:17.661402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:07.373042Z digest=sha256:886d5aa4431d86f1c07a1461b9ce2dae6e8ca03401492d28b3e511898c748b17

Observation 3d05bc31-1160-47d7-85c3-f0b2bef852bc · outbound

This paper cites Scaling Laws for Neural Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Scaling Laws for Neural Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.509003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.509003Z digest=sha256:e770a9f41b5f8e22c3412349f2cd3a2042c51b2a89105d6e0e580e3f64e95f01

Observation 28b1bdb6-0513-417c-a542-9ff30d953169 · outbound

This paper cites Lightgbm: A highly efficient gradient boosting decision tree.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Lightgbm: A highly efficient gradient boosting decision tree

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:17.348865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:07.670501Z digest=sha256:d1a8f2aaf9c3a2b1826d4b7d6c60ac4775f0d6ea21e9c82e593ff448d8130667

Observation fa5eec87-5d99-4c32-adfa-3dd4131dab22 · outbound

This paper cites Looking beyond the surface: A challenge set for reading comprehension over multiple sentences.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Looking beyond the surface: A challenge set for reading comprehension over multiple sentences

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.849579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.849579Z digest=sha256:05c6e1829bf0e0458265fa99cb4eac4224d6377d1ef6bafc59e38b38e65e7041

Observation d515f4e5-b32b-4dbf-80d4-dad73e26adaf · outbound

This paper cites RACE : Large-scale R e A ding comprehension dataset from examinations.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training RACE : Large-scale R e A ding comprehension dataset from examinations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.967559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.967559Z digest=sha256:036c09738a438674f93f6a9d25d5bfa20ffb2ddb46c504cd6de01327443b32cb

Observation 2a9418f5-9f61-4c41-aca4-fbc309ea8ab9 · outbound

This paper cites Not all tokens are what you need for pretraining.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Not all tokens are what you need for pretraining

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:16.966571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:08.109415Z digest=sha256:657813b14cc4a31397a0be3e2ca9f4d1eb9650dbd4860c9913ca6018e9b79b64

Observation 3641c731-e492-4180-a9ec-9cab909a1615 · outbound

This paper cites Logiqa: a challenge dataset for machine reading comprehension with logical reasoning.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Logiqa: a challenge dataset for machine reading comprehension with logical reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:08.255118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:08.255118Z digest=sha256:b58a34b31195ebd298c3ea7feaae5dc33f3ba52f717da50c654a10be2588876d

Observation a96b27b8-a264-4bd2-9990-f3f620b06f1f · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:08.394746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:08.394746Z digest=sha256:9fcfc9dc5877bc7b4166546e66af22ac1979f7fe6b21fa7eebf4b38428a9579d

Observation 15ebcb6b-aab6-4b5b-b14e-3ffaaee57b9a · outbound

This paper cites Decoupled Weight Decay Regularization.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:08.491601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:08.491601Z digest=sha256:723e24c3ba91b6c094a526b573129299f3f415af38428b859e94b0aeb49a5e86

Observation e81c65ec-f9c6-461e-842d-d440cb8af013 · outbound

This paper cites Some methods for classification and analysis of multivariate observations.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Some methods for classification and analysis of multivariate observations

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:16.582732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:08.599517Z digest=sha256:db1bc0809892c98cfa7da430b8e18fb6f6e1643b32632605bb8a516ff9276d17

Observation cda02639-cf84-467b-b204-4ac489126ff3 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:16.337236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:08.801169Z digest=sha256:f53d5909ee78fdb0c545c8aeccab2d1bf8e6f696cfd8b75a1ce1bc130c2b2a54

Observation 3986d171-07d7-4d1c-bec3-6d88314a4849 · outbound

This paper cites GPT-4 Technical Report.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training GPT-4 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:08.901499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:08.901499Z digest=sha256:aed7c634bcbf3b017d81c34d900171b6a1212915526d6e8927e370f5a37c4fb9

Observation e8bb77d4-9e4d-41cc-918e-7953c2e48fe8 · outbound

This paper cites Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fern \'a ndez, R.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fern \'a ndez, R

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.056334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.056334Z digest=sha256:ff59bb10aef97ecb96fe0ba100dbf392a141198ee6cb820b547ff6cab637ea1d

Observation 909bae5a-a467-4000-9335-6ef00d7e9f3e · outbound

This paper cites B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:16.079631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:09.206680Z digest=sha256:2381be80fe1bf10bda30e777312e312f42f7ed9c944f494279b2b10a60fab7b1

Observation 51f31caa-db4d-4702-af24-3e3dcc179a9e · outbound

This paper cites D-cpt law: Domain-specific continual pre-training scaling law for large language models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training D-cpt law: Domain-specific continual pre-training scaling law for large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:15.801841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:09.316683Z digest=sha256:989421bd015c9e993d3c1251275676567780ccf315f2234f248fc6eeb074c42a

Observation df65be68-87cf-4d85-b930-c297b992b76f · outbound

This paper cites Qwen2 Technical Report.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Qwen2 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.397800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.397800Z digest=sha256:071429c1a5028d10ade50527a3509ac0e9502f12ed3d0b4341da373ad053f727

Observation 128ba91b-2acf-4fdf-93c8-924dedff6d98 · outbound

This paper cites Improving language understanding by generative pre-training.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Improving language understanding by generative pre-training

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:15.508809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:09.496631Z digest=sha256:e4b9aad9e103d635b7d66c4c80f46f3d6df49542dc3a154267b04debbd4564af

Observation 7ae6989d-d90e-4cca-80c3-d7965b74c2c7 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.584692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.584692Z digest=sha256:fcf93743c988a0fdabec2cf2cb3737df6fd05bd51644be2d143d7c8608faa756

Observation 9bb47e5a-a264-4656-84c5-30e91f78cbdd · outbound

This paper cites an unresolved cited work.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.654342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.654342Z digest=sha256:0eacf4351a9bfed52a3876632732aa28f791b3dd42336274ba9108e266af9c70

Observation c7c0f0f9-6a3a-437a-9426-cb6192962d62 · outbound

This paper cites W., Hashimoto, T.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training W., Hashimoto, T

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:15.271649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:09.742647Z digest=sha256:314b82ad569f3b23a4cdbd0dec42b994fbb24997b82f405c0eb4fa2b3438a807

Observation 02d10b5b-4148-4fd0-9184-c40bedc80ebd · outbound

This paper cites L., Bhagavatula, C., and Choi, Y.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training L., Bhagavatula, C., and Choi, Y

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.846389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.846389Z digest=sha256:cb45b53dfb24366eaa9179fa75d8fc83a53c87d2e764b0904be74cb34cbdc785

Observation 68496c35-c5de-45b2-9d60-9de0dc544f55 · outbound

This paper cites Social IQ a: Commonsense reasoning about social interactions.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Social IQ a: Commonsense reasoning about social interactions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.950458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.950458Z digest=sha256:0e09bfc01943c377cb1ca65c4eee591e6d1060441402725e515df49aab0a740b

Observation b7c78fb7-fa74-402b-a247-29b0773f37ec · outbound

This paper cites Self-influence guided data reweighting for language model pre-training.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Self-influence guided data reweighting for language model pre-training

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:14.942881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.037663Z digest=sha256:3b97e922b1c8371694c30e9b0a69a2bdc15cda23dc468ee543864a027ea74f25

Observation 453e0082-ae34-4f93-85eb-2bc47ef8a6ce · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.099522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.099522Z digest=sha256:fc667579e3daec4cd6498540f5b6bdd042efe4463b04fb5d477386428bb22436

Observation 567f32ab-526b-40cb-b391-f820d1d0e369 · outbound

This paper cites and Hinton, G.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training and Hinton, G

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.178365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.178365Z digest=sha256:d074a398c30ea60efd74604ffc4c72ceff2b338c2253a5c335a80f5df6e04bb1

Observation f881e0e2-8d5b-4e42-93fc-65d87c471d6c · outbound

This paper cites Learning Dynamics in Continual Pre-Training for Large Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Learning Dynamics in Continual Pre-Training for Large Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:19:11.911632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.282840Z digest=sha256:488320aa693c96f520a9b48df0bf78c560124f8c7ad81724ffa6e7409249888f

Observation 2a633f51-b59d-4224-8915-6a1fad4c7ad1 · outbound

This paper cites RedPajama: an Open Dataset for Training Large Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training RedPajama: an Open Dataset for Training Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.350194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.350194Z digest=sha256:590b5150760573602e47e4075fbe7df6a96648f72c06fb3210dc2551d907c4a7

Observation 94fc3b50-0d5a-4d2e-846a-11952fb1e08a · outbound

This paper cites F., and Gardner, M.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training F., and Gardner, M

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.444975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.444975Z digest=sha256:c6c6576b487d2c5a92c0fedb617668d54071c49ee171ac70854b173857d69221

Observation 5c98e9f7-60a3-4e21-8647-30badaad4d2f · outbound

This paper cites C-pack: Packaged resources to advance general chinese embedding, 2023.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training C-pack: Packaged resources to advance general chinese embedding, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:14.606366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.532719Z digest=sha256:42c80a76dcb0a4c4e3b0d0db182870d7464929c310336dc4a9474a3ddaa676e0

Observation 2f36a394-6086-419d-92fd-cff21e86821e · outbound

This paper cites M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P., Le, Q.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P., Le, Q

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:14.317126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.626223Z digest=sha256:c058bb8213bcfbb5172fcf65946e1393fdc85dd031e8704ee20abd1ae981f32e

Observation 968da2f7-e792-43ea-8804-346f988730db · outbound

This paper cites M., Santurkar, S., Ma, T., and Liang, P.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training M., Santurkar, S., Ma, T., and Liang, P

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:13.932011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.703150Z digest=sha256:d13eafa83b9827c4cb4fa605ad78c8db5500619543c8fcd32962676cbb3eb203

Observation 024a254e-8ef7-4f2c-b59d-0c32fcacb218 · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.793759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.793759Z digest=sha256:6e20426e0cd0273d32199673410230a6cea147989edd04dc63f90adbcabfa643

Observation ed166475-b0e1-41f5-aae6-c00ec04ecbac · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Annual Meeting of the Association for Computational Linguistics, 2019.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Hellaswag: Can a machine really finish your sentence? In Annual Meeting of the Association for Computational Linguistics, 2019

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:13.596553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.914414Z digest=sha256:55e75edcf7db086cbf15141899aa5eed8a888c010c347eaade2c09a81e888787

Observation dbc2900e-8802-4a8e-b9b1-867931298325 · outbound

This paper cites LIMA : Less is more for alignment.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training LIMA : Less is more for alignment

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:13.143966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:19:11.007446Z digest=sha256:535a77cd61570febd751eedad54d9a0b91b1e911cf5b05f9a53a32799ff0b3af

Observation 9e46c7fd-fb5b-4c1a-809a-a2f94e6b0792 · outbound

This paper cites write newline.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training write newline

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:11.124746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:11.124746Z digest=sha256:e5cf90fada1bd7a7e6e292e0a1fbc2bc865f852c9597e09399ed1d482c2d8715

Pith citing papers

Observation 75014b2e-a9cb-4c56-b770-895e9292b063 · inbound

Data Mixing for Large Language Models Pretraining: A Survey and Outlook cites this paper.

Data Mixing for Large Language Models Pretraining: A Survey and Outlook Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.711591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T00:56:04.958757Z digest=sha256:bc8e1ad730cfda54da36011b10e574c73c6ecdba94801a6a7f47ac0210814ee7