Pith. sign in

Paper Citation Record · LEDGER

Data Efficacy for Language Model Training

As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 1 inbound Pith citation observation for arXiv:2506.21545.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21545 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:31:40.416807Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T04:49:28.598430Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:49:52.413967Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3472e36a-d530-47ef-ae56-b2e31262b75a · outbound

This paper cites Training language models to follow instructions with human feedback.

Data Efficacy for Language Model Training Training language models to follow instructions with human feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:19.795365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:19.795365Z digest=sha256:dfee265f107f3159d2639928c7db5b08ed85f5e23ee278d843966143fcdf2eea

Observation 34907434-1239-45b4-88c2-d0f1c4fee130 · outbound

This paper cites GPT-4 Technical Report.

Data Efficacy for Language Model Training GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:34.789743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:34.789743Z digest=sha256:1e04b56edb6454fd06a7b59228dccb942bae36363ccd44c3a3d64d5dfb6dbdba

Observation c20a7caf-e52b-4419-ab88-69ffb3476f62 · outbound

This paper cites The Llama 3 Herd of Models.

Data Efficacy for Language Model Training The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:34.895488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:34.895488Z digest=sha256:40910541e1f5f2baab5ce2832da5b1a84b5eef10c40f627daf1c56dbac5a33ac

Observation 109ce30e-a715-4be8-9bde-903bf3f6e665 · outbound

This paper cites Advances in natural language processing.

Data Efficacy for Language Model Training Advances in natural language processing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:46.230969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:35.076029Z digest=sha256:78d5870315545179e8460e776b3fa68addc06660c2c09a483f56b5c2ae2e73bf

Observation 43b7d785-8f9e-4852-8966-d9661c2e5f0c · outbound

This paper cites Exploring Sentiment Analysis Techniques in Natural Language Processing: A Comprehensive Review.

Data Efficacy for Language Model Training Exploring Sentiment Analysis Techniques in Natural Language Processing: A Comprehensive Review

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:31:40.922314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:35.157846Z digest=sha256:b12d165ba1efd8f10d6af2a8d8d794f297a5d8a787437c463871150b6e7ca122

Observation 13d9fd78-e418-4d90-ab5a-2a45644dba8a · outbound

This paper cites Natural language reasoning, a survey.ACM Computing Surveys, 56(12):1–39, 2024.

Data Efficacy for Language Model Training Natural language reasoning, a survey.ACM Computing Surveys, 56(12):1–39, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.320880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.320880Z digest=sha256:2b04218b84ae545a99e906c41fc04a38496640ca8c4b94783b95d7e8c14a82c3

Observation af796198-10d4-423a-b3d0-5532553f81bc · outbound

This paper cites Ai- based conversational agents: a scoping review from technologies to future directions.

Data Efficacy for Language Model Training Ai- based conversational agents: a scoping review from technologies to future directions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:46.065682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:35.434001Z digest=sha256:4fe2c75b57b56c9d65c9ccf0a95a37ed2590e9c83100520c2fa6e2b1b82f3c38

Observation a1927fc7-0ceb-4778-b125-e5bfacb792d9 · outbound

This paper cites A Survey on Data Selection for Language Models.

Data Efficacy for Language Model Training A Survey on Data Selection for Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.530387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.530387Z digest=sha256:0d11c0114f12ded87f1deacb3466718040973c575610f64ea6ade4f4e61f53fc

Observation ef3463bf-d78b-476a-ac3e-401d1550cb6f · outbound

This paper cites Data selection for language models via importance resampling.

Data Efficacy for Language Model Training Data selection for language models via importance resampling

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:45.895241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:35.679951Z digest=sha256:1ac15626483727b4168ae70e12e92573dd0cb4c8ddc127c6ca139a8e00eb4ed6

Observation e3ec3028-3dba-451f-a437-18ddd3f97bb0 · outbound

This paper cites Data Selection via Optimal Control for Language Models.

Data Efficacy for Language Model Training Data Selection via Optimal Control for Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.732689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.732689Z digest=sha256:a0f209d9a06d8199131292f0d88ae42353dc8151c996bfd0d97f937cd4f74b51

Observation 9afb825a-41b3-4b1a-ab8e-df850b14d8db · outbound

This paper cites Curriculum learning for language modeling.

Data Efficacy for Language Model Training Curriculum learning for language modeling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.794833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.794833Z digest=sha256:199bd7d553546eabf9d45f82ad973ebd950a856dd32d97cca373de9f5b864437

Observation 5455404c-f873-465e-8537-443144bf71ca · outbound

This paper cites A survey on curriculum learning.IEEE transactions on pattern analysis and machine intelligence, 44(9):4555–4576, 2021.

Data Efficacy for Language Model Training A survey on curriculum learning.IEEE transactions on pattern analysis and machine intelligence, 44(9):4555–4576, 2021

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.871785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.871785Z digest=sha256:9056678f057624c1e591b782873044af714c9b6160a4a73c72a12d26f7230743

Observation 181826f8-e99f-458a-b806-4e797e939955 · outbound

This paper cites hello-gpt-4o.

Data Efficacy for Language Model Training hello-gpt-4o

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:45.681972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:36.032874Z digest=sha256:1e498afcec7a2b4bbc0e809c5ae456eba21e4d44b887fbc4a068521a4dad5938

Observation 732c55b2-97bc-4c44-84b0-dcb59c421d78 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Data Efficacy for Language Model Training Gemini: A Family of Highly Capable Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:36.095933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:36.095933Z digest=sha256:08ccc261100ffac43c6d29775917a4f7ab36e3c73302bb4526cc0e839304d659

Observation bfed6da5-5d01-4875-bc9d-f1b78466d95d · outbound

This paper cites Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.

Data Efficacy for Language Model Training Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:36.144310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:36.144310Z digest=sha256:4866804e26f489e6205efc416b32f537d17225302c40317afbd5a018c739ee03

Observation 688f2516-476c-4595-b6e0-f95ff308ca9e · outbound

This paper cites Long short-term memory.

Data Efficacy for Language Model Training Long short-term memory

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:36.249863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:36.249863Z digest=sha256:1af9d2770f34cd4920b9056bbc9bdf58d5d8507cd55fce171a86dd01c43cb9b8

Observation 6d60b150-431e-4a24-a423-184016f8a3e5 · outbound

This paper cites Scaling Laws for Neural Language Models.

Data Efficacy for Language Model Training Scaling Laws for Neural Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:36.341314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:36.341314Z digest=sha256:29fdf83e3d27a14a7f8b11db24a76b1c3f39cbcc0081a9680906e5202ab3ed42

Observation 0858127b-aa58-4d30-8628-25ec4970a281 · outbound

This paper cites Scaling laws for data filtering–data curation cannot be compute agnostic.

Data Efficacy for Language Model Training Scaling laws for data filtering–data curation cannot be compute agnostic

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:45.455618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:36.403199Z digest=sha256:ad8397d2dffe571a72af2a512787a1ad320df02df051761a5788b1d748420d87

Observation 02bebf18-ee54-457b-b079-53b8c8b972d8 · outbound

This paper cites KenLM: Faster and smaller language model queries.

Data Efficacy for Language Model Training KenLM: Faster and smaller language model queries

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:45.203086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:36.509845Z digest=sha256:582009cfa9b356a85f02f1814bd1a533ca630d82e8356b72b03d9aaa5404efd4

Observation 8c87e9eb-75b3-4a0a-9f11-8afdcd28682a · outbound

This paper cites Claude 3 haiku: our fastest model yet.

Data Efficacy for Language Model Training Claude 3 haiku: our fastest model yet

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.984977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:36.629087Z digest=sha256:eee36930a70f8c1b829e1e1f133a94e4134f4e5e225bb3ccd4c6426270c94481

Observation 1252c827-a720-4254-8352-bdd6bde866cf · outbound

This paper cites Paml 4: phylogenetic analysis by maximum likelihood.

Data Efficacy for Language Model Training Paml 4: phylogenetic analysis by maximum likelihood

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.857188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:36.737006Z digest=sha256:fbd52fcd367fedc4afd1f9fc1d2b4e082a641cb9bc7cb2095fdaa311ea8109a6

Observation edcc5ac3-4dc4-4523-a9fe-b72b283f75ac · outbound

This paper cites Common crawl – building an open web-scale crawl using hadoop, 2010.

Data Efficacy for Language Model Training Common crawl – building an open web-scale crawl using hadoop, 2010

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.730683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:36.842436Z digest=sha256:e24150221bb64d7dea8ad5ab3eb9adf511500a58c33240123f7eb930feeffb44

Observation 2a3c4282-7a05-45db-aa0a-e57051d3bb3b · outbound

This paper cites Project gutenberg, 2004.

Data Efficacy for Language Model Training Project gutenberg, 2004

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.626745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:36.925600Z digest=sha256:95a34aff9dde0b54cf16514bef90c6390b84062acbc9acbd2b2b21b3290a2468

Observation 1800c862-3d77-4e66-8e4d-b2ff96e97ab8 · outbound

This paper cites Synthetic data for deep learning , volume 174.

Data Efficacy for Language Model Training Synthetic data for deep learning , volume 174

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.504460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:36.983223Z digest=sha256:4bbf5453db81f2cce79cf823cb6f7e2cfee63708cb7008d2a135bfa83097828c

Observation e1b7d25e-5cca-4c8c-8ae9-ab772e92dc82 · outbound

This paper cites Virtual sensors: Abstracting data from physical sensors.

Data Efficacy for Language Model Training Virtual sensors: Abstracting data from physical sensors

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.406796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:37.055947Z digest=sha256:52e9ecf7116e90950f19ac0728c70c453ba5a50c06aa2b2e78ca1fac87934c4e

Observation ef5ba448-e607-4b88-95f1-db06e6d7af76 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Data Efficacy for Language Model Training Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.178879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.178879Z digest=sha256:f48299278833e39c1ba5e7c6142572252fd138d1f928accd01a61d553ba53ea6

Observation 53478cd4-9bc4-4938-9b6e-df65e110288a · outbound

This paper cites The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only.

Data Efficacy for Language Model Training The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.320180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:37.243725Z digest=sha256:da9f4e2f359507b8f3e06eb930c8eabc4f17dd86c7bad07e22ad346c4e66bbee

Observation 61cb4c60-bc85-4f79-8533-cfae6e3eb1ff · outbound

This paper cites Redpajama: an open dataset for training large language models.

Data Efficacy for Language Model Training Redpajama: an open dataset for training large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.336209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.336209Z digest=sha256:ab546946ef924aa74794abc76d69f10b0bcc78d171dd710dfa1189a969add304

Observation a9af43c7-a9b7-451e-b517-6d0514dd2b55 · outbound

This paper cites RedStone: Curating General, Code, Math, and QA Data for Large Language Models.

Data Efficacy for Language Model Training RedStone: Curating General, Code, Math, and QA Data for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.409614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.409614Z digest=sha256:7ebc955597f516cbc6f40a8a32a800488c958756297433309564f7de3db68a4b

Observation a2af7a3f-b91f-433b-9166-60a72cc70ff1 · outbound

This paper cites Mates: Model-aware data selection for efficient pretraining with data influence models.

Data Efficacy for Language Model Training Mates: Model-aware data selection for efficient pretraining with data influence models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.202949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:37.516009Z digest=sha256:4f3f806827ef5c447b655b0185a4139ebad0cb2dbd4ab7c29bfc1ae593d52e91

Observation e7f4349a-1d10-4a6f-a369-fef8925df5bf · outbound

This paper cites Training-free dataset pruning for instance segmentation.

Data Efficacy for Language Model Training Training-free dataset pruning for instance segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.095050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:37.572752Z digest=sha256:0fd79f377fbdf8a2e979c2647454021515771a7d9928b5f4ad46981238ad0f6d

Observation 5ad266e4-2031-4862-899b-a22d6be16703 · outbound

This paper cites P-diff+: Improving learning classifier with noisy labels by noisy negative learning loss.

Data Efficacy for Language Model Training P-diff+: Improving learning classifier with noisy labels by noisy negative learning loss

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.990600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:37.620112Z digest=sha256:12b844dc1fb1b390ed23c656eaa26aa50e32a8296913f1cebb72336c50f7858c

Observation 3fb4ae5a-4f96-48e0-9d3f-a804118be1dd · outbound

This paper cites P-diff: Learning classifier with noisy labels based on probability difference distributions.

Data Efficacy for Language Model Training P-diff: Learning classifier with noisy labels based on probability difference distributions

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.881483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:37.665766Z digest=sha256:386c1e7bcf85907cb2ad142719b14ed4d6bbe172db9c865fbfa8afa55e92b02a

Observation ea0935cd-49b8-421c-b4b8-883d969337d8 · outbound

This paper cites SemDeDup: Data-efficient learning at web-scale through semantic deduplication.

Data Efficacy for Language Model Training SemDeDup: Data-efficient learning at web-scale through semantic deduplication

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.703358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.703358Z digest=sha256:ba639f97a735aad14e1620ecd72ca0f2be8c8f39e245d31ac4ffe7e835d1ff16

Observation e655be23-c353-4bf2-8cc8-0f06301c35cc · outbound

This paper cites D4: Improving llm pretraining via document de-duplication and diversification.

Data Efficacy for Language Model Training D4: Improving llm pretraining via document de-duplication and diversification

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.768535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:37.798234Z digest=sha256:dfab88b144bd433d6b410b16fca70d006c0f25e8d60615a1538c566169a3e9e3

Observation 13b16161-46e5-45cc-b23a-cb4b4c4a5f00 · outbound

This paper cites Strategic Data Ordering: Enhancing Large Language Model Performance through Curriculum Learning.

Data Efficacy for Language Model Training Strategic Data Ordering: Enhancing Large Language Model Performance through Curriculum Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.853648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.853648Z digest=sha256:b3923491bfa7a5532264b3cc5a196029699ebdaacdcdb0542249fcfb9cf885dc

Observation 4b295c33-2616-43ae-a09a-71f2730342b9 · outbound

This paper cites Does the Order of Training Samples Matter? Improving Neural Data-to-Text Generation with Curriculum Learning.

Data Efficacy for Language Model Training Does the Order of Training Samples Matter? Improving Neural Data-to-Text Generation with Curriculum Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:31:40.710309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:37.971943Z digest=sha256:4960558126479f26bfaea1889bb737fcc8a5e0768603a92ad27f11626fd1ab27

Observation 8c0be7cb-f617-4dd3-8ec5-df298404a13f · outbound

This paper cites DoReMi: Optimizing data mixtures speeds up language model pretraining.

Data Efficacy for Language Model Training DoReMi: Optimizing data mixtures speeds up language model pretraining

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.673371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:38.143000Z digest=sha256:483fc66224d35e3ae24f33895aa0156822b6259bf16f64b6b3b2ca064e962ded

Observation 0b2a68e7-4a68-4024-9700-cf5ce660c34e · outbound

This paper cites Lima: Less is more for alignment.

Data Efficacy for Language Model Training Lima: Less is more for alignment

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.563723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:38.258542Z digest=sha256:031a265601faa86cd439e63b769533969189adc0ffdb1ec0a1026defac6fdd49

Observation ca58f8fa-b5bd-49ec-963f-0d6d9132a0b2 · outbound

This paper cites Openwebmath: An open dataset of high-quality mathematical web text, 2023.

Data Efficacy for Language Model Training Openwebmath: An open dataset of high-quality mathematical web text, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.370698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.370698Z digest=sha256:f4ae87a374d763dbbdc1be1519bff533b8c66713edd5bc11ad0e30534f022689

Observation 3eb84262-2755-4f47-9c16-039dbd257e40 · outbound

This paper cites MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics.

Data Efficacy for Language Model Training MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.509344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.509344Z digest=sha256:03685b0edb1f17df46d5682389d4c77ec57606bb720fa8d430ee166b89bb6350

Observation b18f4e11-793c-4db5-b71e-5e668be63f4a · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

Data Efficacy for Language Model Training StarCoder 2 and The Stack v2: The Next Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.679922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.679922Z digest=sha256:e1af66321ca9e8afb674a398e201343a1d97dff390e8090c9f6b7f6d7b2faa9d

Observation f1c0a048-7337-4760-b423-606723280f2d · outbound

This paper cites Epicoder: Encompassing diversity and complexity in code generation.

Data Efficacy for Language Model Training Epicoder: Encompassing diversity and complexity in code generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.794928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.794928Z digest=sha256:dc759242d68ca22fb191af5ebcc7a3e9f6b1215e1788c4676339fa1e80104bb6

Observation 7eaf93f6-b84b-47a9-af78-a170358f2de2 · outbound

This paper cites Mistral 7B.

Data Efficacy for Language Model Training Mistral 7B

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.889964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.889964Z digest=sha256:91c2d74a06ec917426cbf7baf48bfd5db70e0fd631956f162715db7063499359

Observation e5ee4398-ee2c-41cb-a1c9-f5e6ec0c7bf8 · outbound

This paper cites Qwen Technical Report.

Data Efficacy for Language Model Training Qwen Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.997852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.997852Z digest=sha256:15038ca8071656e17be9be5855411297b9d43e1abaab3070ec4294dd6a82448c

Observation 3bb581e0-945e-4cf0-bbb4-8478a8f754fd · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

Data Efficacy for Language Model Training OLMo: Accelerating the Science of Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:39.060536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.060536Z digest=sha256:15efc7bdb0d71dc4baf9e1ff1ee2bf061057614b37ab52dbb3863c408ebc5811

Observation 0e32dec9-e955-44f6-b0b1-f9a89ceda0eb · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of ACL, 2019.

Data Efficacy for Language Model Training Hellaswag: Can a machine really finish your sentence? In Proceedings of ACL, 2019

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.416290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:39.121670Z digest=sha256:83934862a6f47b6aab9ce56daacf65ad0f2912d6f240fa0d52113442d3cb0942

Observation d30f3ec4-a3b1-498c-b6e9-edfd1ef90b55 · outbound

This paper cites The winograd schema challenge.

Data Efficacy for Language Model Training The winograd schema challenge

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.234992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:39.177098Z digest=sha256:2748055c24e2b5f69555cbd2aacd633e208f314795aaf45443dea16a884cb149

Observation 1f834b77-8d9d-473d-b3fb-7b758a0a26b8 · outbound

This paper cites The lambada dataset: Word prediction requiring a broad discourse context.

Data Efficacy for Language Model Training The lambada dataset: Word prediction requiring a broad discourse context

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.085204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:39.220870Z digest=sha256:7c2851631e46b628077a4b2fbbb2d33c2e2dc2a22f975af2233a1a87eac01c28

Observation 6aa43bff-0821-4ecb-9854-5b3586e92bc1 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Data Efficacy for Language Model Training Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.927964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:39.290035Z digest=sha256:7a1d72e10afdbc2dac6e4d884fb2bfdc6bf186f702ea0fc660f9271f837bdc3a

Observation a029a61f-bf3a-4d63-8253-7a5d63fb8916 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Data Efficacy for Language Model Training Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:39.363914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.363914Z digest=sha256:4dba7125305b259a0dc856fb8bb153bdba66b5ab779219d7e541b840a2b06f4d

Observation fc872b53-e93a-4a05-a77d-65bd721bfd29 · outbound

This paper cites Piqa: Reasoning about physical common- sense in natural language.

Data Efficacy for Language Model Training Piqa: Reasoning about physical common- sense in natural language

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.710576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:39.444920Z digest=sha256:281e1c0e2ff83f057ca7b8ff1a532060ec0f26b23f2076e49b7858ff14b73378

Observation f407e16b-dd00-4152-aab2-32f538352ad5 · outbound

This paper cites Liu, and Matt Gardner.

Data Efficacy for Language Model Training Liu, and Matt Gardner

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.527336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:39.518996Z digest=sha256:4b3b0752b025b6baf7d1b6677e7538c8ffd30bd82d597a49d8f7364c7c5efbea

Observation 451ce2a6-f0a4-45f6-9ee0-4059911f2b09 · outbound

This paper cites BoolQ: Exploring the surprising difficulty of natural yes/no questions.

Data Efficacy for Language Model Training BoolQ: Exploring the surprising difficulty of natural yes/no questions

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.323259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:39.599850Z digest=sha256:c4dc854af7d7b6e08bd17f32d91556cfc2d9a05002c980cccd4c44ad6fd6354d

Observation fa2ec6e9-33ea-4869-8e1b-9ea6d509a64e · outbound

This paper cites an unresolved cited work.

Data Efficacy for Language Model Training Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:39.689642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.689642Z digest=sha256:fc1c8a57ff8806d6bb1f0f0b681df46bef24a194a6a54213a45896fe14668ce2

Observation 17350254-b8c4-48a4-8322-bfa482950b43 · outbound

This paper cites Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019.

Data Efficacy for Language Model Training Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:39.764554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.764554Z digest=sha256:ab761c836d68a5d39a8747f07579fe4366867b3d0afb7db923ff76cb9eeda393

Observation 9caa4a74-9795-4b26-ae42-fdb2e2fb3e54 · outbound

This paper cites an unresolved cited work.

Data Efficacy for Language Model Training Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:31:42.162493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:39.849319Z digest=sha256:81e6e39de07e0e6495bada2a919b045ab67d3a3b215ee55ce1472dc8382693f1

Observation ed73ddf4-2c92-400d-a2e8-1fc17282ccd3 · outbound

This paper cites Program Synthesis with Large Language Models.

Data Efficacy for Language Model Training Program Synthesis with Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:39.955344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.955344Z digest=sha256:9e983aa4a9794538315639c36f33dd00fc62103b688fe9a73a52e3e4c51cd6fb

Observation 752b0d3d-0c92-47dd-aad2-1212e2992466 · outbound

This paper cites Efficient large scale language modeling with mixtures of experts.

Data Efficacy for Language Model Training Efficient large scale language modeling with mixtures of experts

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.973753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:40.039407Z digest=sha256:967829617186fb41ead460e9f7f24e53492bb01ca14ea52ef43ca697d0b220c4

Observation bd642587-8300-4094-999f-0cd460e8c61d · outbound

This paper cites Decoupled weight decay regularization.

Data Efficacy for Language Model Training Decoupled weight decay regularization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.786740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:40.140435Z digest=sha256:d7188f036e41a335474720ac4f29d5a73b7f53afa9db6c61b2be096b105fca8c

Observation 2bc119d4-5638-4b74-be28-6607582b63c6 · outbound

This paper cites Comparing the pearson and spearman correla- tion coefficients across distributions and sample sizes: A tutorial using simulations and empirical data.

Data Efficacy for Language Model Training Comparing the pearson and spearman correla- tion coefficients across distributions and sample sizes: A tutorial using simulations and empirical data

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.578495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:40.240414Z digest=sha256:5b4dc8633d1ab35a90779a347059dad0241369a18fc1fd99fd68748b62f9e14c

Observation 6d65bff9-d365-460a-af49-54a81ef4cd17 · outbound

This paper cites As described in algorithm 1, we apply a linear transformation to the mean-pooled representations of instances along the sequence length.

Data Efficacy for Language Model Training As described in algorithm 1, we apply a linear transformation to the mean-pooled representations of instances along the sequence length

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.369377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:40.328843Z digest=sha256:b208295bbbaf200cdb6f6e762131a6cfe8b60ec4c11fb513e1ff724658027901

Observation dc04d5f7-4d10-422b-bf68-e651558f3866 · outbound

This paper cites I am overpowered by the discovery of my own genius for management.

Data Efficacy for Language Model Training I am overpowered by the discovery of my own genius for management

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.149390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:31:40.416807Z digest=sha256:6d05c4b3d69218a6eacd5fc29b1d4ed550059a18ff995fb5af0901d5807037c9

Pith citing papers

Observation e2a8aeb6-0d71-470b-aa72-6e671db70eb7 · inbound

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning cites this paper.

GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning Data Efficacy for Language Model Training

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:49:52.415420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T04:49:28.598430Z digest=sha256:de9771492848386165361433a7e6d8250ff7470881e9709f0a8bd89383b70e88