Pith. sign in

Paper Citation Record · LEDGER

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

As of 8 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 22 inbound Pith citation observations for arXiv:2506.20920.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20920 v1

Coverage vector

measured 100 of 118 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:46:42.375506Z

measured 122 of 122 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:16:45.263479Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 118 outbound references displayed

  • verified exact8
  • verified fuzzy2
  • unresolved89
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5ec22fa5-a012-41c0-82da-f189097500a9 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Yi: Open Foundation Models by 01.AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.435145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.435145Z digest=sha256:f62b496b6e4ad891ff8d9e3f8d46d5c48dd16a5732a07fc3e7467730ae3878ed

Observation 93645589-9dc1-4d2a-8128-c1a94669f570 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.489072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.489072Z digest=sha256:a20ff26f01f042c84ea23c2a5cb8ee91b3397e1860745fb44b332fb3d814aea3

Observation 6894d93a-aa44-48df-8279-75e4b1f5b8f6 · outbound

This paper cites Llama 3 model card.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Llama 3 model card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.557062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.557062Z digest=sha256:f32f43c61da49112dabf17d8bd46869f7af87578087ea742b29317456a07ebd0

Observation 9b2b22c6-7c9f-43bf-b997-1ec8c9115267 · outbound

This paper cites A survey on data selection for language models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A survey on data selection for language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.625972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.625972Z digest=sha256:466487aab88485dcd4f9ca4a640b727e8280d75c2d55f2b525f5fddb5c9cfdd9

Observation 767dab6a-c430-4bea-a8fa-31aa9e767aec · outbound

This paper cites Open llm turkish leaderboard v0.2.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Open llm turkish leaderboard v0.2

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.702254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.702254Z digest=sha256:4e8f6d50832a0875bbb1a81e7af3634a0c5a9e7e04878d1a0059fed1c4dd9bf1

Observation a1f05eab-e33f-40fb-977c-3aac94923944 · outbound

This paper cites A l G hafa evaluation benchmark for A rabic language models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A l G hafa evaluation benchmark for A rabic language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.797026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.797026Z digest=sha256:23a2e789b1d3e6646a102724fc15d346dbd876cb8c318e952e2f5c996fc8bf8c

Observation c00582aa-05a4-47a7-96da-83d765789020 · outbound

This paper cites 101 billion arabic words dataset, 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language 101 billion arabic words dataset, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.843819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.843819Z digest=sha256:837ee65cc98bb26d2623783506f2eec446a4f94e97000c291734dc7e87096877

Observation 20cce32f-fab4-43bc-8d2e-76412bc7fc45 · outbound

This paper cites On the cross-lingual transferability of monolingual representations.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language On the cross-lingual transferability of monolingual representations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.881227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.881227Z digest=sha256:79e2505cdd4a22536cebb1f36d2a6d77a890b5f1df3aca0ad59e1750fb4a7e57

Observation 67eaca72-07bb-4902-b2db-fa169d3747a0 · outbound

This paper cites A call for more rigor in unsupervised cross-lingual learning.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A call for more rigor in unsupervised cross-lingual learning

Reference 9

Resolution
verified exact
doi, observed 2026-08-06T22:46:43.869917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:33.948108Z digest=sha256:04a5a268b53181cea8babda276cae7bc7040b75330f84b36d018b0ca61ead2e1

Observation fd5889bf-fa8d-4af2-aa2b-5ae7d4a83975 · outbound

This paper cites Japanese massive multitask language understanding benchmark, 2023.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Japanese massive multitask language understanding benchmark, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.995458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.995458Z digest=sha256:062eee2fd84286b9df589e5e7d5ebdb5810f9a400761238e6c824f5aab32fdc0

Observation 9b0d82cf-0d66-40f5-8a0e-da42bf5d4814 · outbound

This paper cites The belebele benchmark: a parallel reading comprehension dataset in 122 language variants.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The belebele benchmark: a parallel reading comprehension dataset in 122 language variants

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.075955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.075955Z digest=sha256:56493e9a9dca442400d3967f04e5c02f0a904dcd6b3953799b77cd621edbfbcb

Observation 2d29a952-39ae-463d-87d5-d28ef38c1316 · outbound

This paper cites Building Machine Translation Systems for the Next Thousand Languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Building Machine Translation Systems for the Next Thousand Languages

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.150087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.150087Z digest=sha256:2bb108829c4c23ff0adcfe01d072c8fd323c4c5b263f6f0b460201ee79fb22a2

Observation aee1a052-c635-4462-b70b-a26749d0af24 · outbound

This paper cites Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.210423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.210423Z digest=sha256:57b1b871c7ec30a96aaaa93ad7af112dffd26821b7db1b42b72abf5fa688ced5

Observation c3578c5c-d4c7-49a8-a653-2ca0dfa9d86e · outbound

This paper cites On the resemblance and containment of documents.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language On the resemblance and containment of documents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.319292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.319292Z digest=sha256:a6174e230adcbaf20de2f7829938cd7513071489e753aaeb9b095fe8bf0966b3

Observation ead3e28e-4569-40f6-878e-666442a5ef75 · outbound

This paper cites An open dataset and model for language identification.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language An open dataset and model for language identification

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.474853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.474853Z digest=sha256:921e589abf4117dee19538ca36959fce9d7e2613cc7ca514ee969dcbed3d7884

Observation 6f6c3cb5-c38d-4cce-b6f4-f6f4b634ed5a · outbound

This paper cites An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT).

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.583082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.583082Z digest=sha256:9bf6720f128ccf189edc2c9a264f2fcbcbc6b706062e4dee3672fe0e3e3fb5ee

Observation 51c0b978-de55-4141-810a-32a5f05bf112 · outbound

This paper cites PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.736671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.736671Z digest=sha256:91d9a48b9865e927c407a3f033c26cdc7b38702c0091324d8e76317e672b074d

Observation d97dbaa2-3c1b-4e30-bb53-124de1eeb969 · outbound

This paper cites Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.831748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.831748Z digest=sha256:20537693c3a3ae33bee16f2226876527455bfefa2fa0c329456b996c680025eb

Observation f48cd5d1-bcbe-49cb-8ec9-074466a7393e · outbound

This paper cites Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki

Reference 19

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:46:34.988316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.988316Z digest=sha256:f344fef72370dd4a9a525ec792ee003777d963fbf312370975d589e8b63d78a9

Observation a73efc9f-d303-4a64-a261-dd8ea894f2dd · outbound

This paper cites Command r+.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Command r+

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.082235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.082235Z digest=sha256:f53580e887d08057df6065f647cb61ce7bcd11a9960036db468809e241a5da6d

Observation 204dfc02-b923-4ba9-85cb-9b4525f83964 · outbound

This paper cites Unsupervised cross-lingual representation learning at scale.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Unsupervised cross-lingual representation learning at scale

Reference 21

Resolution
verified exact
doi, observed 2026-08-06T22:46:43.742969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:35.190771Z digest=sha256:e995fdd13c9b67832e48549ce8860611f590196cf214410e51ae85238ce800be

Observation ebfebb77-afd4-4762-8ae6-1ec465b2ef68 · outbound

This paper cites Neural learning for question answering in italian.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Neural learning for question answering in italian

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.333775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.333775Z digest=sha256:93f3e9974a739b46a81d5b5b3cc96b20a3cb1106bfe1447f0735e88dd436412d

Observation df62415c-318d-40c7-80cc-884a41aa911f · outbound

This paper cites Dataset for the First Evaluation on Chinese Machine Reading Comprehension.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Dataset for the First Evaluation on Chinese Machine Reading Comprehension

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:46:45.068840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:35.445116Z digest=sha256:5dbabe1c7c31204f22765612af39a179688205d1e8ac270acc82a00c8f057319

Observation 504fbcb6-49c0-4587-9919-3ba7ae8ddc77 · outbound

This paper cites Daniels and William Bright (eds.).

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Daniels and William Bright (eds.)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.543260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.543260Z digest=sha256:bc660104dcb14aa006760fa0e25971b6e37f0407d65ce7cc5df8b693da84fa6d

Observation 75a59b1b-7199-4d01-ba28-fac54f1e3df1 · outbound

This paper cites A New Massive Multilingual Dataset for High-Performance Language Technologies.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A New Massive Multilingual Dataset for High-Performance Language Technologies

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:46:44.888181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:35.727398Z digest=sha256:e62aef6dd00845ca05dac4f7455dcafef293b5823fe88a1372ce316b58859613

Observation ed081d84-e88a-4f48-92f3-42000a1ef7c8 · outbound

This paper cites BERTje: A Dutch BERT Model.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language BERTje: A Dutch BERT Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.817846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.817846Z digest=sha256:32d1215bd923464e31100246b0ab276864b439ab29b51698ccd7971d3eeb31ad

Observation 488d3324-dc9d-45c4-a85d-3f57bcfdbcf5 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.894130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.894130Z digest=sha256:b5084e3328be34c106a0f212f3ec55648642e024a9d99b8f5c11b0c7a985cada

Observation 579cc685-a676-4f08-aa87-8aeca05345ef · outbound

This paper cites RobBERT: a Dutch RoBERTa-based Language Model.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language RobBERT: a Dutch RoBERTa-based Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.985633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.985633Z digest=sha256:805471cf11c4412b0cf47d981075534943c774a9fe13870821ae1327c30fd8e2

Observation 48891c63-e204-485a-b9db-06320ce3391e · outbound

This paper cites FQuAD: French Question Answering Dataset.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language FQuAD: French Question Answering Dataset

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:46:44.741197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:36.078826Z digest=sha256:55336e826aeaa7b6a92773933ceabdfb66aa396da7c769764cb84f02bf90af32

Observation bf5d4220-13a8-4211-ae7c-eab29ffc1d12 · outbound

This paper cites Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.172908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.172908Z digest=sha256:b813520854aae94b380953a9dff985cdf9a45e34523e1333650df83d814f948d

Observation 96fee737-29a8-477b-b0ed-85d088c6a9fc · outbound

This paper cites Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.260283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.260283Z digest=sha256:df1f04fce5f2d41b2544b127e6d6e0ac0f22bf1a2bae40d6056e21d9f81fee9d

Observation c3edfc52-178b-4997-99d4-9ae23425345f · outbound

This paper cites Eberhard, Gary F.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Eberhard, Gary F

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.357550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.357550Z digest=sha256:37686876deaccb2bfb724094c9800a79d4169a6af365bc7e1887ed1e82f42ea2

Observation 086ea4a7-24ab-48bc-a1e0-83a70fe57691 · outbound

This paper cites SberQuAD – Russian Reading Comprehension Dataset: Description and Analysis, pp.\ 3–15.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language SberQuAD – Russian Reading Comprehension Dataset: Description and Analysis, pp.\ 3–15

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.484014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.484014Z digest=sha256:35e411db4e26600e3935696489b1666cb1375597ad0082a4ad5ce39bc2a1d317

Observation 4f6ba979-b408-4848-b114-102f4590aae4 · outbound

This paper cites Arabicweb24: Creating a high quality arabic web-only pre-training dataset, 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Arabicweb24: Creating a high quality arabic web-only pre-training dataset, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.563221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.563221Z digest=sha256:6e48ff7eded992691f423701301cc7c18bed52a2a979f1b52208f1ddce8a372a

Observation 44f6f449-cb49-424c-803c-bad6e90723f9 · outbound

This paper cites Guerreiro, António Loison, Duarte M.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Guerreiro, António Loison, Duarte M

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.638830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.638830Z digest=sha256:3130856c74d39fad88a806d7c99ff6a0a9eba6b41f24a2b5613e1f5739eabd7e

Observation be88cc1e-59c7-4bed-bca5-0a003eddfb9e · outbound

This paper cites MERA: A Comprehensive LLM Evaluation in Russian.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language MERA: A Comprehensive LLM Evaluation in Russian

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.710892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.710892Z digest=sha256:e9b889c27ca212e994e509e07dc6b74e8cf60479f7454feeb5e81c11aae56ba8

Observation 4fd9627c-c59c-4bd2-9299-0f896ee29e4e · outbound

This paper cites Open llm leaderboard v2.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Open llm leaderboard v2

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.843320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.843320Z digest=sha256:90a6bd9a2fca4e06b59fae811a7560c36cfd3de332aa771520f73dfc88fad9f1

Observation fb661be8-d0a6-4bce-a8a6-859bc397a9bf · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Gemma: Open Models Based on Gemini Research and Technology

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.956706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.956706Z digest=sha256:982f1c4b1989ee43acb0904b84437cf1dda5d13242ec68d9d2be9653e7072c19

Observation 981672ec-3d2a-4080-855f-8cb4f19284c7 · outbound

This paper cites The Llama 3 Herd of Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The Llama 3 Herd of Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.083695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.083695Z digest=sha256:63add6026ba2cd5255ee18fdf5fc5054a9f1b39935e54a6b1e5e8b9ef4bc3ba0

Observation 94c3acd5-d761-4e6b-9598-f18da8e52e00 · outbound

This paper cites Studying Large Language Model Generalization with Influence Functions.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Studying Large Language Model Generalization with Influence Functions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.222530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.222530Z digest=sha256:ec60743c3beb3dc5e51991539b0be413f6159bb595e98f8cb3359f93df39f8e6

Observation 7072c36e-25f9-4ad5-9295-6e4aa128b307 · outbound

This paper cites OLMES: A Standard for Language Model Evaluations.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language OLMES: A Standard for Language Model Evaluations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.365791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.365791Z digest=sha256:92efc5bbb9e4809c86c228292b63e2e0800da01addf41f88a57705eca4cf2941

Observation 9a33ca50-0ec2-4830-a8ab-8355fa9c5c2c · outbound

This paper cites EXAMS: A Multi-Subject High School Examinations Dataset for Cross-Lingual and Multilingual Question Answering.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language EXAMS: A Multi-Subject High School Examinations Dataset for Cross-Lingual and Multilingual Question Answering

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:46:44.529872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:37.521847Z digest=sha256:e91023b7cf8de6eef86a6291768e1a7194fcb189efcc2c8cd771c5759d5f2844

Observation 1d7babeb-2c21-4c95-8506-9831e9a7cab5 · outbound

This paper cites Measuring massive multitask language understanding.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Measuring massive multitask language understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.613762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.613762Z digest=sha256:ee5c527d982d3f6b48be1996b9f3c5d0dde30e61562a5f17c16adaed1f016e7c

Observation e6cfbcb3-9399-4019-bdd0-090d0eb25307 · outbound

This paper cites Khmer natural language processing tookit.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Khmer natural language processing tookit

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.677450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.677450Z digest=sha256:dfef168cf52ad92cbdd321763ec4655c163025b760702086a5bb33fa95414acd

Observation 59ff659f-ad21-44bc-867b-0135918901f7 · outbound

This paper cites spaCy: Industrial-strength Natural Language Processing in Python.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language spaCy: Industrial-strength Natural Language Processing in Python

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.821687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.821687Z digest=sha256:9752eb91c8633a4fd1126fb484a31b2e012ead4c2f5b29b2e90b134e2a295b6d

Observation 561fdd31-a08e-42b2-bf8f-c9cb73492142 · outbound

This paper cites C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.938316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.938316Z digest=sha256:5e107012cb6eb2665c2523d97e955b066e8ed14aa8cc98745e9b0cd1bb4d8f3d

Observation ac51f328-e38f-40e5-8db0-f5462da4da08 · outbound

This paper cites Mistral 7B.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mistral 7B

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.104108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.104108Z digest=sha256:949ab9821146cf413bf0abe6ccc821165fe4f1059e7145154b32b45bdab2166d

Observation f4ff6bd0-52bb-4dfb-9c64-fbd5f186fe61 · outbound

This paper cites Mixtral of Experts.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mixtral of Experts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.254676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.254676Z digest=sha256:f1d3b9b6f1cad91d403312b1165509d9509c7184629ead575da622e84eb83b70

Observation 1bff3b4a-7fb8-4da2-8b86-0d413dfd1696 · outbound

This paper cites The state and fate of linguistic diversity and inclusion in the NLP world.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The state and fate of linguistic diversity and inclusion in the NLP world

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.411289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.411289Z digest=sha256:af6a80b37010a39508591c669e8c6389ff956ffb3bfa73feec9f2330659ab145

Observation bf778262-2509-4a1f-8333-9248fed2b284 · outbound

This paper cites FastText.zip: Compressing text classification models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language FastText.zip: Compressing text classification models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.542891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.542891Z digest=sha256:fd729e6c6532507565fd450eee891dffc7448aa06213d2daaefc9db1d8a324de

Observation 0d33314e-4afd-42de-954f-636bcf10589b · outbound

This paper cites G lot LID : Language identification for low-resource languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language G lot LID : Language identification for low-resource languages

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.664452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.664452Z digest=sha256:e38109915c2834c8c07c740dc01931f6bb6a584456ea43bbf30b634d38571217

Observation f91fd40f-c527-4989-b2c1-955c2585825c · outbound

This paper cites Glot CC : An open broad-coverage commoncrawl corpus and pipeline for minority languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Glot CC : An open broad-coverage commoncrawl corpus and pipeline for minority languages

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.729060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.729060Z digest=sha256:479cfcf843494315de93e6c5b58193ee796b022d4b75d1cdf9fa0673f8342872

Observation ab275403-84c5-447e-86d0-96b6802e314c · outbound

This paper cites IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.868060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.868060Z digest=sha256:c30499d736d458793cc08e2a583b5942b6b6185dc7acdeb254c5164cabcff7fb

Observation bf5468c1-6200-469f-b7b6-e5dccd617073 · outbound

This paper cites Large language models only pass primary school exams in I ndonesia: A comprehensive test on I ndo MMLU.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Large language models only pass primary school exams in I ndonesia: A comprehensive test on I ndo MMLU

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.953519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.953519Z digest=sha256:d95fd4e10a50ed01e95f18a8c40281cf68a30e42acd19400a321dc7a725fee7e

Observation 7a77dfd9-3344-4fa0-ad00-4b36543393b8 · outbound

This paper cites ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.027217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.027217Z digest=sha256:136198537f7db2460230c6d171e364a969a9e3e636e06abc50f1fbf03cd2dea7

Observation 96434442-1e71-403d-96da-6cea0a979b97 · outbound

This paper cites The IndicNLP Library.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The IndicNLP Library

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.099607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.099607Z digest=sha256:b8bddee6cf0b52753db650ab54d7d2847dacaff7efb324fb3bfa97d457af0efc

Observation aecd3d84-6991-4604-9374-a45c68a30b04 · outbound

This paper cites JGLUE : J apanese general language understanding evaluation.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language JGLUE : J apanese general language understanding evaluation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.169580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.169580Z digest=sha256:f12a07aad2d9a0bec4e9250a441c4f47d8b8f14a9eac915dae3660ca7eb3d8b5

Observation 8ee3b8ed-aa0d-41e0-91be-f4186e87631a · outbound

This paper cites Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.239094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.239094Z digest=sha256:515c849a707a739901f7224ac259594dd909c5aa85630bc65496147de0b539d8

Observation eaeafdc2-d5d4-495e-8f1b-8115b58b28c5 · outbound

This paper cites F lau BERT : Unsupervised language model pre-training for F rench.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language F lau BERT : Unsupervised language model pre-training for F rench

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.343322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.343322Z digest=sha256:854f0e7e34484b58cb5124c57e218c039db9bc4bb0af03359c20cd8bee33ca72

Observation 957c7581-a3b7-4084-a9cc-2950589ca17d · outbound

This paper cites Open-arabic-llm-leaderboard-v1.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Open-arabic-llm-leaderboard-v1

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.454235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.454235Z digest=sha256:a9d722ee7662af1c638047ef15f2115f9214074ad8dc7f16cf7fe9568a41ceaa

Observation 98be8b18-486a-4cbf-ac8a-6ea4421f473f · outbound

This paper cites Deduplicating Training Data Makes Language Models Better.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Deduplicating Training Data Makes Language Models Better

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.550459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.550459Z digest=sha256:a4f96b59eda5826ab2128acae19a6d8cedc765ff067b7c2e8c3cf2423d8dd68d

Observation b6b25d1c-e7e3-4731-87d8-d9ad2d2f05ab · outbound

This paper cites Kiwipiepy: Kiwi package for python, 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Kiwipiepy: Kiwi package for python, 2024

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.641312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.641312Z digest=sha256:83807c8c5c993e63484e7b275d5fdd1c49403c7e1c951b1d68fa0796bb05cf76

Observation 072cbc6c-d94e-493c-9e08-bb7b96b8a997 · outbound

This paper cites MLQA: Evaluating Cross-lingual Extractive Question Answering.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language MLQA: Evaluating Cross-lingual Extractive Question Answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.718750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.718750Z digest=sha256:e6a78cb49596ab018154ba32983247c5579d6b208462b2ba7461d5d84aacd269

Observation 82379ae4-9e75-4a1f-b422-34a70798f7d2 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language CMMLU: Measuring massive multitask language understanding in Chinese

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.789508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.789508Z digest=sha256:025bb3d4e01466c66004e6a29c0a290ddeaacd049b36ccb25e976e3921c28e36

Observation c6e28839-9292-470b-9c28-2b6c7cf4ae35 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language DataComp-LM: In search of the next generation of training sets for language models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.887384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.887384Z digest=sha256:41c01d195796b655e3add2ee8e7958d1844abebf017abe635c9753464c7ff4f2

Observation d4208a9a-d1bd-463b-bb92-2efd48be185b · outbound

This paper cites Common sense beyond E nglish: Evaluating and improving multilingual language models for commonsense reasoning.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Common sense beyond E nglish: Evaluating and improving multilingual language models for commonsense reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.950982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.950982Z digest=sha256:1e86f7573e8b30c6e8c62305d2fd36d1097360e89101e58697077d781290ca50

Observation a01ec591-bec5-493d-bd80-54c62c0db555 · outbound

This paper cites Few-shot Learning with Multilingual Language Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Few-shot Learning with Multilingual Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.120607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.120607Z digest=sha256:4478aa6f3e0ebe95959b454aef86e329bfaf7ec25762f8aa5ae8cf4bf5bf8347

Observation adc5788a-159a-4393-8e78-f5702c2bec77 · outbound

This paper cites FinGPT: Large Generative Models for a Small Language.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language FinGPT: Large Generative Models for a Small Language

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.228960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.228960Z digest=sha256:5f88820f0e09f5650210d45e4dfee5733fe1a253f0a65c9c4e0b2b598e5efdd7

Observation ffa8195f-6c42-430e-a18d-f5c38ae41b92 · outbound

This paper cites Quantifying Variance in Evaluation Benchmarks.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Quantifying Variance in Evaluation Benchmarks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.311443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.311443Z digest=sha256:633d0d6cd1b061ff3e79837ad0c61bf874980e9281aeefa1fa380d45262c59c5

Observation 4be25e36-faac-4254-9da6-e62f2ffad73a · outbound

This paper cites C amem BERT : a tasty F rench language model.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language C amem BERT : a tasty F rench language model

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.376511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.376511Z digest=sha256:f9ab94c8f587ce2d2de1a988f8fb0071fe99f5b778db724e339b70ffdcde15fb

Observation 30ff154b-ccfc-425f-815b-7b6da12380b0 · outbound

This paper cites Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.446794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.446794Z digest=sha256:ac80f34d71da18dd723127391d8f8db99e17145684ef7490a7291232a02aaaae

Observation f2027824-04d2-46e6-8767-f458b66c54f0 · outbound

This paper cites Mnbvc: Massive never-ending bt vast chinese corpus.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mnbvc: Massive never-ending bt vast chinese corpus

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.530378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.530378Z digest=sha256:53f890856a1d4c5872cc8059e0264e4e94eb725c44c18aa5f13bce57f12b3373

Observation 996645bd-7901-40f4-a1f0-8f1ca508ca5d · outbound

This paper cites Neural A rabic question answering.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Neural A rabic question answering

Reference 75

Resolution
verified exact
doi, observed 2026-08-06T22:46:43.644012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:40.612521Z digest=sha256:4b4733fbdd9065b130a783e4937f91cbad4d71d70d52f25772464edae875a066

Observation 8441a717-8366-4f55-8bd0-2c7e762cf8be · outbound

This paper cites Crosslingual generalization through multitask finetuning, 2022.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Crosslingual generalization through multitask finetuning, 2022

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.708439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.708439Z digest=sha256:3d90d3ba85dc3b126541ffe51bc82768af8e0233c1532648b4f8b334e193201b

Observation f82ef352-d1ca-4cac-a91d-b5d765dd15d9 · outbound

This paper cites Rossi, and Thien Huu Nguyen.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Rossi, and Thien Huu Nguyen

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.789698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.789698Z digest=sha256:d69f204fff1fe7d76a97c29d408320717221f711342859596466b53b06ef0064

Observation a4a84145-1fcf-45ba-930a-8d8230b3334f · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.884821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.884821Z digest=sha256:14d5c1f8748b08557d4e8318da6e8622b10e3857404f3451bf248840c3a867ce

Observation b02b7449-4b62-4a97-bf29-1576201f80f5 · outbound

This paper cites Omnia russica.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Omnia russica

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.940992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.940992Z digest=sha256:65ca06bae0a3390c40067bb1bb149a9b2480fa73bab89727788194b9c2b1ec82

Observation 967d7d78-2977-47f4-8d9b-6bc072bb1a8d · outbound

This paper cites Botok: State-of-the-art tokenizers for tibetan language, 2025.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Botok: State-of-the-art tokenizers for tibetan language, 2025

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.003899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.003899Z digest=sha256:c5c036c2357b8fe96ce295ff47d9d68d8fa653529ab18ab1d269968e6eeebfc7

Observation 41d85cdc-552c-43ca-8bde-c85e358c797f · outbound

This paper cites Building pre-train llm dataset for the indic languages: A case study on hindi.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Building pre-train llm dataset for the indic languages: A case study on hindi

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.104136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.104136Z digest=sha256:29a2e6d7648d8688a07893950bf0683886478d7f36471f871783446ed2f9cef6

Observation 13ff38a0-299a-4cff-8429-f354bf0330b1 · outbound

This paper cites Hellaswag-th, 2023.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Hellaswag-th, 2023

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.162030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.162030Z digest=sha256:a4f4c4dbf4c38e4a434892aee1238fbebc92512f00c40637e6cfb04058241f90

Observation 3636b3d9-ad11-4cf7-aa9b-70f4aad44a20 · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.208321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.208321Z digest=sha256:d9e05e4f77eac483725089f7ebeb5fdcfd70e2e3d70ab69050b6b32e5c3bd5c6

Observation 1421b3f6-cc9a-4d8e-9540-d4986d14c6f1 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The fineweb datasets: Decanting the web for the finest text data at scale

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.292670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.292670Z digest=sha256:9b431f88a10dcdb0d6fa575ec0579e9c7b58175450911168444b2187ec7cd964

Observation f6c30583-bdce-4a57-86ea-01fa20c397b2 · outbound

This paper cites Laonlp: Lao language natural language processing, July 2022.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Laonlp: Lao language natural language processing, July 2022

Reference 85

Resolution
verified exact
doi, observed 2026-08-06T22:46:43.524600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:41.328021Z digest=sha256:0badc12084cb447718462b2c308811706392875c84b36a2fafd7e3225856fba0

Observation d4910a25-953d-42a0-a5cc-a60dfbd80149 · outbound

This paper cites P y T hai NLP : T hai natural language processing in P ython, June 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language P y T hai NLP : T hai natural language processing in P ython, June 2024

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.368981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.368981Z digest=sha256:5ddd0e08bd1b97becd611f7af5cbdd6f58b977f045d6f62a74ecd13fe78e3d0b

Observation 792e8fff-e19d-49a1-8874-bc9c2e996040 · outbound

This paper cites Typhoon: Thai Large Language Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Typhoon: Thai Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.432866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.432866Z digest=sha256:306d077e4bab35d79bdcd16e6234142d08c51ccd70642ce5c7271109ff61970b

Observation 00ab1a9a-a5b7-47ee-a3ee-117d04220b57 · outbound

This paper cites Pllum: A family of polish large language models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Pllum: A family of polish large language models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.515692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.515692Z digest=sha256:f63402cb7208a11784d379dd916235a8d8c559ae1f18fb75906e2a1790dff418

Observation 800b7eb2-19b6-4ae1-bbfa-70d87904a68d · outbound

This paper cites Chinesesquad.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Chinesesquad

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.584393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.584393Z digest=sha256:333f3873ba13ef7b5d733cb0627f99a8a2bc0bd39bcd31b23370b138476b307c

Observation 8ae140c9-8575-42bd-a403-2f666001c830 · outbound

This paper cites XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.630958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.630958Z digest=sha256:fd9acd9d5e8ebc286b1617eb611ba9394c3f8d5b962a37b71677b0bfe538a7cd

Observation 213d288c-62ad-464c-a708-530c35a4b8d7 · outbound

This paper cites Stanza: A Python Natural Language Processing Toolkit for Many Human Languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Stanza: A Python Natural Language Processing Toolkit for Many Human Languages

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.693007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.693007Z digest=sha256:54a0ae9a54bbf6399f55ae34fc2742a85aa96840c831310fba5c511d6ef4fdbf

Observation 16b60db9-7c02-4cb9-90ab-2fdbfc3e98e6 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.766456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.766456Z digest=sha256:d19ea47e83ead2c8ddde822ae70383417274660d889ade7e07ceb2fdfb7d2a93

Observation b23d45d5-4bb6-4328-afa5-fa2979fd8043 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.821123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.821123Z digest=sha256:5fb09adc7965c6f5ab6d6aa77f9bba819dab369abc09481d7efeb0ab985131e8

Observation 078640bb-dca3-4481-b332-812e91aec898 · outbound

This paper cites Impact of Pretraining Term Frequencies on Few-Shot Reasoning.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Impact of Pretraining Term Frequencies on Few-Shot Reasoning

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.895184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.895184Z digest=sha256:64c5ca31d0495eba229f8f156c4c72d107012fc2b388032975e84ffc967a75c0

Observation f195de96-efca-4b7d-8ce1-43649ab4f262 · outbound

This paper cites How Much Knowledge Can You Pack Into the Parameters of a Language Model?.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language How Much Knowledge Can You Pack Into the Parameters of a Language Model?

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.932613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.932613Z digest=sha256:a7cdcec2d12a440826c8414984b58d68b78a1daa962b2ee60407df92de5b384b

Observation e9647ceb-cb4a-488a-a35e-0db91dae8727 · outbound

This paper cites How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.999731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.999731Z digest=sha256:1b3e5852860a025a7786c0d388b305c09577b4b53a8a22fa4cdb02a18bf11075

Observation 4b72f10a-e44f-434e-9e6c-32f17e657e5c · outbound

This paper cites Pyidaungsu: Python library for myanmar language, 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Pyidaungsu: Python library for myanmar language, 2024

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:46:45.888138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:42.080031Z digest=sha256:109e6fd6caca530b95835ce5d9e3decc1781e338a1fef3a333052c700d7fa06b

Observation 63725564-5a04-489c-a39b-9f21105c3bc5 · outbound

This paper cites Compact language detector v3.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Compact language detector v3

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:46:45.879187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:42.103931Z digest=sha256:65beceff13364fe0f62b43e128d48184cba0db785089335a81427db83477177f

Observation 8f76eddb-8eb6-4b46-a375-2dc8251a7cbd · outbound

This paper cites Mintaka: A Complex, Natural, and Multilingual Dataset for End-to-End Question Answering.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mintaka: A Complex, Natural, and Multilingual Dataset for End-to-End Question Answering

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:42.192719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:42.192719Z digest=sha256:c8054649b5cabf27ee8dd5072399a8cf5d6f938a9b26b364165aeda42dddc264

Observation 6f7c332a-cbf0-48f2-8dd8-067bbe507e12 · outbound

This paper cites INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:42.248951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:42.248951Z digest=sha256:82f7ea3e002d95f1552e46171bdb912b9683e98b216962f61d10e8d5fe1b1c3e

Observation 982ab6e6-25d6-4030-8a74-e6e7278fcfaa · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:42.309351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:42.309351Z digest=sha256:76a002e07adfbc9fbd407c2b634b6ea1dc4fb98a9993e1e18ed3ae3de40434c8

Observation c20a06fc-aaf4-4213-a34f-b98e46657df4 · outbound

This paper cites Thquad: Turkish historic question answering dataset for reading comprehension.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Thquad: Turkish historic question answering dataset for reading comprehension

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:42.375506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:42.375506Z digest=sha256:0f44e90cf2db35927633ed5c8450a77e209fbc252a1f41286a6d9bb1ec997697

Pith citing papers

Observation 09e337de-4ea6-40da-aabe-c6305c9f5365 · inbound

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation cites this paper.

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:45.263479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:16:45.263479Z digest=sha256:0c9a6a8c2cba03d37335cb8b0b70c186a99b106ad872f60117f08cd437c8dd38

Observation 0c6bb196-3a1f-4b67-ae83-0f7f9f57eabe · inbound

Observation of momentum dependent charge density wave gap in EuTe4 cites this paper.

Observation of momentum dependent charge density wave gap in EuTe4 FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:45:15.379631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:45:15.379631Z digest=sha256:996f055ab41da32e5e09856405ba4acae7cfc65166482578b08c1994cb1ccbab

Observation 67cc670d-1c49-4e0f-9050-c5fdcc28d126 · inbound

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models cites this paper.

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:50:08.479194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T17:50:08.399160Z digest=sha256:7f3a8190ee3f671576eef1feb9d6d3a08ace488052c5d9f8e1ccbc8647b8f5ae

Observation 7b73d5d6-df5b-429e-b414-072673213de2 · inbound

Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations cites this paper.

Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T05:42:50.078953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:42:50.078953Z digest=sha256:3828833cf07d6de0d89ba89b9e4d3db58085fcca733a6ae817d1d5d19cc20373

Observation 45b04929-9776-4355-bbde-389fd154ecf0 · inbound

The Effect of Scripts and Formats on LLM Numeracy cites this paper.

The Effect of Scripts and Formats on LLM Numeracy FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T08:58:50.692729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:58:50.692729Z digest=sha256:19d042562e66f6f12ed74f802e05c48bf79dd5d7918ed8c51c68bcbc61de8a27

Observation e263f088-3f4e-4306-8aae-d48749d9a534 · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 167

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:27.454474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:27.454474Z digest=sha256:9a393affeac6f346786c88fbbf295c658b48326edeb710ca6fd13e9d008f5520

Observation 79a8722b-9b1e-435f-9f17-7bc8c9414d93 · inbound

OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report cites this paper.

OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T23:37:59.719719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:37:59.719719Z digest=sha256:e299284a1d7e98b4067fe6d1d2e47dbe82e2d47f3d2542dde0a46ba934dbb8ed

Observation 763e30af-8044-4a58-8526-9f21ffcb16c1 · inbound

Scaling Laws for Mixture Pretraining Under Data Constraints cites this paper.

Scaling Laws for Mixture Pretraining Under Data Constraints FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:48:00.949304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T21:44:31.429223Z digest=sha256:f97a43f8dd9d4a69e7c2f72d8de65597e1365dd9a8b2f6f1fb0cb185bcd23a40

Observation 3a3d4225-0be4-4675-bc7f-70b881a3c0a0 · inbound

Scaling Laws for Mixture Pretraining Under Data Constraints cites this paper.

Scaling Laws for Mixture Pretraining Under Data Constraints FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.604606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T16:36:30.007014Z digest=sha256:b34e07a3cb7085faacbe12e712139724ad3ea0d4d49dc2c47626e7caab56f4b1

Observation 10a7a53c-b215-4103-8017-0c99cb34c20e · inbound

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings cites this paper.

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:27.448259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:17:26.661595Z digest=sha256:4d9ab763779316cf298693327925cc5db9113428f6c52e0b285f6aafeb3dd501

Observation 009b4acd-a370-4dba-a981-20466c2442fe · inbound

Granite Embedding Multilingual R2 Models cites this paper.

Granite Embedding Multilingual R2 Models FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:17:34.932496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T18:16:49.148303Z digest=sha256:3ef01a9430e236c81e1b47aa74aa3e94ee55970e887798c28e5f84b0ec508d08

Observation 9c65265e-e2be-427e-a18f-c02d71c8ba7e · inbound

A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE cites this paper.

A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:13:13.616472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T11:09:22.027588Z digest=sha256:923e5d77761b41fab53f6ecc1da4dd55ab50d4b4ce5dd7767f669e5d00de973b

Observation 55dab8fb-b24c-476f-a89e-a007b018028b · inbound

SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark cites this paper.

SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:38:12.634240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T10:34:49.783942Z digest=sha256:c5ee9b00580be6dd4c6d33f1669898dadc2b7b8a6645704c4c902f5d85ebdac3

Observation ed65c45e-2831-430a-8d11-502aa052c297 · inbound

Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions cites this paper.

Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:15:20.257197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T04:11:06.312503Z digest=sha256:66741ffafc4bfe053c0d6e12e203eb489f2b152c54359e56aef0c7cb6823c191

Observation cf0129b7-a60b-40ca-9fb4-83adcd620143 · inbound

Mimir: Large-scale Multilingual Concept Modeling cites this paper.

Mimir: Large-scale Multilingual Concept Modeling FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:24:38.681356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T11:08:44.027943Z digest=sha256:1a8de6a9cfd85ad202ba8f779000e495f58af03774b0062ef3d47a78e5095c1f

Observation 57f65719-6f44-406a-861c-fcf0faa9aabd · inbound

Ishigaki-IDS: An Open-Weight Verifier-Aware Model for Information Delivery Specification Drafting in Building Information Modeling cites this paper.

Ishigaki-IDS: An Open-Weight Verifier-Aware Model for Information Delivery Specification Drafting in Building Information Modeling FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 25

Resolution
malformed identifier
arxiv_id, observed 2026-07-02T23:17:29.057127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T18:24:53.592902Z digest=sha256:4cf6b02aa0f501c6137bff0fe457da8aafb42cf19d575d12715ddd2e311a23e8

Observation df6a9db1-36aa-41d6-b236-56f1861d6a01 · inbound

Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention cites this paper.

Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:29.893255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T17:40:06.036019Z digest=sha256:317568f64ef75515c993efef25e1861ce662e9fd93d8cba6acef1e152ad6f5d2

Observation d2bd763f-f994-4834-9947-29e54e4eb04e · inbound

moBERTo: A Modern Encoder for Portuguese via Continued Pretraining of ModernBERT cites this paper.

moBERTo: A Modern Encoder for Portuguese via Continued Pretraining of ModernBERT FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:19:44.124861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T10:07:55.521194Z digest=sha256:7371108035b1ec3462ddbd8d608450aedfb857df7804e126549e7c10f81bfb28

Observation 1b5b1bdb-58fb-408b-9c29-2bf74907bb15 · inbound

LangMAP: A Language-Adaptive Approach to Tokenization cites this paper.

LangMAP: A Language-Adaptive Approach to Tokenization FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:49:46.653379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T08:26:33.185340Z digest=sha256:e0a119f90cd466bb5e1b71ce7b19fa9abff44af395f3d1e37a90a098c317d65e

Observation d3e45960-722e-4f56-bb55-437ffa4ab47a · inbound

SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment cites this paper.

SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:00:09.022875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T20:48:04.003465Z digest=sha256:c3e3863a322c8a7526e45baaa0f6461b2a90f68e3d9b91fcc8ed55b6389586ff

Observation 65447489-0aee-4602-976e-e5023ff38cb3 · inbound

MultiHashFormer: Hash-based Generative Language Models cites this paper.

MultiHashFormer: Hash-based Generative Language Models FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T04:23:05.508816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T04:13:05.082903Z digest=sha256:945d17929c8efaa03f2c7bea5306aebc902256dd104e1c939af3f4997b2d1739

Observation a12afaeb-34ce-473a-91f9-9ad1b79e3171 · inbound

In-Place Tokenizer Expansion for Pre-trained LLMs cites this paper.

In-Place Tokenizer Expansion for Pre-trained LLMs FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:16.706422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:16.706422Z digest=sha256:c75aba64d87b8a038661798e9ae1c564cc2a986dde47575f5ac3bf28a5417734