Pith. sign in

Paper Citation Record · LEDGER

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs

As of 7 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 2 inbound Pith citation observations for arXiv:2507.03253.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03253 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:21:57.501475Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T07:01:44.900291Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T05:33:58.451084Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact3
  • verified fuzzy17
  • unresolved46
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7450d016-0969-4188-afac-599e12bcdb68 · outbound

This paper cites GPT-4 Technical Report.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:49.221314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:49.221314Z digest=sha256:140ea4de6cd2a16d33cb0885c930b6e279e1d094be6276758b472a1edc88dfc0

Observation 79574b7c-16fa-4f40-8ca6-671d9b6310ec · outbound

This paper cites an unresolved cited work.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:22:01.871394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:49.280948Z digest=sha256:f6e051a94af373b5be004be20beb8545d8a287083a9e2d9c82256e1f1ad3b3fa

Observation 8fa02c83-c199-41cc-879a-1af021779bf1 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs The claude 3 model family: Opus, sonnet, haiku

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:01.684652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:49.382321Z digest=sha256:4d2ac950829a76d909d170047d90f9192ce441220b6a7f9f55ee9b05f53ff914

Observation d44f2571-9608-44ad-848e-a9abfb514dac · outbound

This paper cites Program Synthesis with Large Language Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Program Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:49.508104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:49.508104Z digest=sha256:5038df65235e66c7f9fc18333ce0c98d5e69b31eb563eb892d91fae53216c862

Observation 2a620327-9de0-4ad6-8008-923297ebb377 · outbound

This paper cites A critical analysis of the largest source for generative ai training data: Common crawl.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs A critical analysis of the largest source for generative ai training data: Common crawl

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:01.574642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:49.662920Z digest=sha256:758341b4155ca558576339c7a8363840b09d089f1b30b8501bb1f2de93b9cfff

Observation 53b6e27e-e9fe-4b74-b5df-c96e9cf58780 · outbound

This paper cites Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:49.842658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:49.842658Z digest=sha256:98c980b90a25f9c4a31baad82006d03195444f9329f5ff35d6bc1225254ecd14

Observation 90a0aa22-8055-4f12-8749-da15131acabe · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Piqa: Reasoning about physical commonsense in natural language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:49.980601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:49.980601Z digest=sha256:c13ac8895ecbb594815d8e54d9ccd2383c0563a6dd14614a5956d6cd8007074f

Observation 7f531adb-7629-4bee-b65c-538d2d47463b · outbound

This paper cites On the resemblance and containment of documents.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs On the resemblance and containment of documents

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:01.424074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:50.161331Z digest=sha256:e14415aae485e8854abc6f932d557f49c76013089b3588ce75086685bdc194af

Observation f50b8dbe-fa86-489f-8281-d08ba7434a6b · outbound

This paper cites Evaluating Large Language Models Trained on Code.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Evaluating Large Language Models Trained on Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:50.343916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:50.343916Z digest=sha256:8c3936de6f518d196943dfdb2ba936d0501e9496e4fa751b227f610d1cddc9cf

Observation 8c03c560-6257-464d-a824-1f44df122b5a · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:50.502340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:50.502340Z digest=sha256:241f7ea07271306a7a048e6c3bf61e7d8163663bb555a98f13673384b140b8b1

Observation 1f456238-d246-4437-b4f1-387d691d94ce · outbound

This paper cites Flash A ttention-2: Faster attention with better parallelism and work partitioning.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Flash A ttention-2: Faster attention with better parallelism and work partitioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:50.621007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:50.621007Z digest=sha256:5902f80c7cb798cefd8bcd1322f51c03faa10f7beb8d4ac5ce19773ff998f292

Observation ca42f1cc-ecfb-413a-946a-75339c94f4d8 · outbound

This paper cites Sailor: Open Language Models for South-East Asia.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Sailor: Open Language Models for South-East Asia

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:21:57.715855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:50.740610Z digest=sha256:08c325d6f5bf06279d5927fba3ef74a7ebfb856159e97be865d68ca331f025cb

Observation 00a45189-c2c1-48c3-98c3-a66350e3491d · outbound

This paper cites The Llama 3 Herd of Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:50.905266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:50.905266Z digest=sha256:68f350b0d9831f63037b514720d3a98b08a09198d9a0c64397941d2aebb97721

Observation b117b13d-bb9a-4808-8711-8260ab07d018 · outbound

This paper cites Minedojo: Building open-ended embodied agents with internet-scale knowledge.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Minedojo: Building open-ended embodied agents with internet-scale knowledge

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:01.267272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:51.070802Z digest=sha256:1eba6f5f07100c9304cfb1083fd1df4ee135a033527c5f85b2cef9b1a5aee104

Observation 7dc05f53-411e-40c0-bacc-be4be248d23d · outbound

This paper cites Lighteval: A lightweight framework for llm evaluation, 2023.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Lighteval: A lightweight framework for llm evaluation, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:01.120985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:51.224440Z digest=sha256:fac370fcd21c5fe19356978e0117c21b862de4ebd6fa80fa0073a00e5e741dea

Observation 940cbd60-b228-40c5-aa52-27d22bb6a829 · outbound

This paper cites InCoder: A Generative Model for Code Infilling and Synthesis.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs InCoder: A Generative Model for Code Infilling and Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:51.342397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:51.342397Z digest=sha256:5ab95f81a6caa76d7b14fcd5087175582741da73e7a973ed708734e22ddbeef1

Observation d0b5aa86-d09c-4d75-bbd7-699cf929be26 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Measuring mathematical problem solving with the math dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:51.486728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:51.486728Z digest=sha256:330ba061a753be0a7e54372e24811702cdd8875577c6e1361cf7160745941d9f

Observation b30ca208-ed28-44c2-82af-2ae001040afb · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions, 2023.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:00.978871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:51.629900Z digest=sha256:29f2e6206f0f449f3480c8efb843daae337edc8b458b8ce45aadedad46b7ad45

Observation fc1bd634-0ab6-49a8-9a90-94e64049b0bb · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Gonzalez, Hao Zhang, and Ion Stoica

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:51.719872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:51.719872Z digest=sha256:d58c5a1c345f8971d79dfe55c51ca58b924f22eaefd6fbb418c28df87b42f364

Observation be8874e2-b0a5-42da-9f9e-cffd68cd07d8 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Rho-1: Not All Tokens Are What You Need

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:51.821393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:51.821393Z digest=sha256:0ff509162a2550de88c4323e8d5d22aa8d67b0f5542bae78015d962df5b9bcff

Observation 165cd0f0-ee77-4662-8d89-b2ee17be86e4 · outbound

This paper cites Efficient Inference for Large Reasoning Models: A Survey.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Efficient Inference for Large Reasoning Models: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:51.911641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:51.911641Z digest=sha256:b392fc7a90468da9030f966bd0de897a99a04f146469d24e5e2de5f6f1423c7a

Observation 176c9530-0d86-42cd-81b3-64832d10da35 · outbound

This paper cites OpenELM: An Efficient Language Model Family with Open Training and Inference Framework.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs OpenELM: An Efficient Language Model Family with Open Training and Inference Framework

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:51.990657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:51.990657Z digest=sha256:30681c435f2b745d017d7a2c7d966d198dc978ce7d9e3f1311811ec33a3f649b

Observation 8bebc2c8-8413-4360-9d74-d46b248dce4f · outbound

This paper cites SLANG: New Concept Comprehension of Large Language Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs SLANG: New Concept Comprehension of Large Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:21:58.461364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:52.082429Z digest=sha256:fc47370a71d3239f7ca98bdff8ef48fb89f65184077ca18c23958f4d59a9e8f5

Observation f6c9158a-dbfe-4a3a-9b99-50336ebd57b9 · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date, 2024.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Introducing meta llama 3: The most capable openly available llm to date, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:00.766841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:52.176927Z digest=sha256:e81586cbe710c012c15b145d561a4b0ce7f075b8ac252cca88ce96d20fe0e5ba

Observation 9ae6e40a-e297-4276-b15c-456dbc3e33c2 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:52.245731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:52.245731Z digest=sha256:71ce8e64f647aa7cbd741a86cbf26b2e5ca1b62edc863f7e363b594c3770d464

Observation aa198cee-3b96-4822-b4c5-75c112b9d7d1 · outbound

This paper cites Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:52.353944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:52.353944Z digest=sha256:efe3fd5907204fcf31ce9792f281c884a605e7115a867ebf07ad6d78b1656916

Observation 5b978836-5c46-4da4-a3f6-08f365f46dad · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Generative agents: Interactive simulacra of human behavior

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:52.474419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:52.474419Z digest=sha256:71a06dd3b960d3df0b5adfda3bed93e42ce490a4ed8308569039542fee318d14

Observation 3be0ad10-5659-45d7-85d6-078a19e0f096 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:52.628204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:52.628204Z digest=sha256:09bc627200e4ac65a17707a75336a60f98655f5acf4566ddee21e6892801de65

Observation 3b214c50-b851-4640-a7d8-d2f1d12bb48b · outbound

This paper cites Datatrove: large scale data processing, 2024 b.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Datatrove: large scale data processing, 2024 b

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:00.595823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:52.744819Z digest=sha256:ba3460ffec8d4de1d199820300b5a7509e9ce6a2715d89d6247ded89867647c0

Observation ef8dcdae-7b02-4ae4-b247-3dc043503be2 · outbound

This paper cites The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:00.412211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:52.879117Z digest=sha256:a0eabc1a336e22d113eab2a6cd9833aedf73a3b75ec6b566d021b4c09ea44f64

Observation e2972caf-9d02-45b9-b472-b5c476447db3 · outbound

This paper cites DataMan: Data Manager for Pre-training Large Language Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs DataMan: Data Manager for Pre-training Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:53.021981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:53.021981Z digest=sha256:dfa94f95bb4a2b91f015102ae91b34b12849c652801407a71f595c217602fdd6

Observation 2b88ea0e-48fe-48ed-a7a7-911089c81335 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:53.148940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:53.148940Z digest=sha256:0a6f167221594043855ef6f14f0edc2c21b2c19f3c21c66053b942903f7565f0

Observation df44e4de-4cbc-40f9-944c-b146bd7f49f8 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:53.276242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:53.276242Z digest=sha256:5421746995e8c3a35018e32857eee4c8cf722655c8fd809aec860724c6792239

Observation 9babaecc-3f96-47d9-8570-8ce6e469ea9c · outbound

This paper cites Web data mining with organized contents using naive bayes algorithm.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Web data mining with organized contents using naive bayes algorithm

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:00.139969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:53.406768Z digest=sha256:859ce8ac20ef4bb274c7aa17ef16392497e0fe6343f3b4ca6b163beef6df29b7

Observation 4fea4167-8318-4f55-8c65-2fc992a212fc · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Winogrande: An adversarial winograd schema challenge at scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:53.567288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:53.567288Z digest=sha256:ab9a3dbb350c33ea72df8a195365a381529b7ec7edbc77182e1c07a13e2feb48

Observation de97d225-d551-42fc-9daf-9ec83628ca63 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs SocialIQA: Commonsense Reasoning about Social Interactions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:53.691167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:53.691167Z digest=sha256:08a7a1865d7df0b2401f956dc50de076fccaa2276f874d853c18bc36244f4a63

Observation 75c4317c-ccea-4174-a1ff-5c9aa60e5590 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Toolformer: Language models can teach themselves to use tools

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:00.002087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:53.844678Z digest=sha256:f264c5de002e41454566f321e977ebe24a3dcb6d3e4031c724122e1b98bcd54e

Observation 9621e435-1282-4009-87e8-92e5d1623dc4 · outbound

This paper cites Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:53.998507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:53.998507Z digest=sha256:2416dda567577fff8afd9c3477a11612fbfcd213630457155fe364c63945e70a

Observation 4e0e0e2c-b9bb-4f69-897c-f218fddc8c6f · outbound

This paper cites SlimPajama: A 627B token cleaned and deduplicated version of RedPajama , June 2023.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs SlimPajama: A 627B token cleaned and deduplicated version of RedPajama , June 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:21:59.833764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:54.173260Z digest=sha256:380b6cfe82f665907755dc2b2f913be822fb4f5a832e25454699c5346cf2e5a2

Observation f9a3b0d5-5a17-4bd4-a6da-0334819f2b0c · outbound

This paper cites Dolma: an open corpus of three trillion tokens for language model pretraining research.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Dolma: an open corpus of three trillion tokens for language model pretraining research

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:21:59.545747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:54.325158Z digest=sha256:2da6a5deedaff1fa2a9f2bced1446089edfa4a5cfd5c2b91584da724e4560dce

Observation 34c86beb-14da-4524-bff9-e9b876b99569 · outbound

This paper cites Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:54.439779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:54.439779Z digest=sha256:9033589cb914ee511891a1c20d9a287dd22cc8121223b47274935a3ab1bf7e77

Observation 3e3606bc-1499-4c96-8dd5-121609bd9e60 · outbound

This paper cites C ommonsense QA : A question answering challenge targeting commonsense knowledge.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs C ommonsense QA : A question answering challenge targeting commonsense knowledge

Reference 42

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:21:54.624141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:54.624141Z digest=sha256:8a257e73c83c68d7c4f7076caafe7ec03b6762f3830ec53abd9d8ce406b13df4

Observation 5b1017bf-c110-4a24-9483-2bccce464167 · outbound

This paper cites Redpajama: an open dataset for training large language models, October 2023.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Redpajama: an open dataset for training large language models, October 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:21:59.401589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:54.781545Z digest=sha256:35dcb046b3b17070dc8f42876e1a597645912ecad5b54b2ccd015c924e697d18

Observation 0e9793d1-3fec-4858-a7bf-134458de5c9b · outbound

This paper cites M Towhidul Islam Tonmoy, S M Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs M Towhidul Islam Tonmoy, S M Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:21:59.219332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:54.951767Z digest=sha256:6cc953baecaaecdbc32db15388eb6b2371b4f0069aa3df13bc3adf705e97cea5

Observation 8738f8af-7e15-4c73-8233-0a045088789e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs LLaMA: Open and Efficient Foundation Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:55.088543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:55.088543Z digest=sha256:ef1d2a44412cae981808d94f27ccaeae69ae4c79986f9e6c0e77605756271f9b

Observation 5d90a520-0610-4b87-af93-cbc33ab5fd3e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:55.236194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:55.236194Z digest=sha256:f0e1de938b9660905d25feb133e06991ce9ed0a5213e9434244573df1b42bb4f

Observation c47027db-4378-4261-9fd5-3943afbc1431 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Code Llama: Open Foundation Models for Code

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:55.411568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:55.411568Z digest=sha256:0a74db975d63dc0115613b76f48d9a079c5e517f0be2964c1f4aac45f7a1ed47

Observation 41e22c1e-de88-4907-9578-997b904eb7bf · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Chain-of-thought prompting elicits reasoning in large language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:55.555671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:55.555671Z digest=sha256:e0f11554ebd8fe10f4443bb0d3a22f7682a181bbbafd0fcdb5629635b9d87945

Observation 7f417b6c-ba9a-4484-8601-42c4466c8828 · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Crowdsourcing Multiple Choice Science Questions

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:55.665550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:55.665550Z digest=sha256:c9fa05ba2d935fb27ddc34f7e50b485ea9efa15b1b4d3e8eb8dafd19107f6588

Observation 594b4f42-1bf6-46a6-9157-bd1f6a7078d3 · outbound

This paper cites QuRating : Selecting high-quality data for training language models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs QuRating : Selecting high-quality data for training language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:55.821166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:55.821166Z digest=sha256:7ad6b1961aa6213459e887013dba684701bb25aa9fdf66fe11d93082586f6069

Observation c19a7082-0c5e-4986-893e-1ed92657c4fb · outbound

This paper cites Data selection for language models via importance resampling.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Data selection for language models via importance resampling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:55.957694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:55.957694Z digest=sha256:c837d1667a307e58b1dc58c0da532989b974672abd4fb4152de76052b1c79f29

Observation 692d263e-4ffd-4a57-a9c3-087be751aa98 · outbound

This paper cites Qwen3 Technical Report.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Qwen3 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:56.118676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:56.118676Z digest=sha256:3e7a0ea177fc83830161eb2edcf6e9fa42798020f374331e3f8e6f1a47d28794

Observation 8d2a5181-c345-4c75-a198-2b59d2c5aafd · outbound

This paper cites React: Synergizing reasoning and acting in language models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs React: Synergizing reasoning and acting in language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:21:58.985198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:56.290379Z digest=sha256:0990edd981bd58da5926c28812be734670076b01b2996b99d3218163227a3dae

Observation 591e6031-5101-4b14-9a15-3885759666b3 · outbound

This paper cites Craw4LLM: Efficient Web Crawling for LLM Pretraining.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Craw4LLM: Efficient Web Crawling for LLM Pretraining

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:21:58.169963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:56.432140Z digest=sha256:559899a32569747241e3e8f1b2738f5abf521ad19e7c211dde415aac402aa3d1

Observation f3e0ccfe-6668-4d09-b5f5-7a27ce0e6b6c · outbound

This paper cites MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:56.521904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:56.521904Z digest=sha256:3bc348ee4b73160f4873d25bd0fbd473e907aa77a051c6e4ed0e7fd3f82d6bd5

Observation ef9f0231-9922-4d4c-b621-f86de036139a · outbound

This paper cites A normalized levenshtein distance metric.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs A normalized levenshtein distance metric

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:21:58.795332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:56.621387Z digest=sha256:dedb14d9727115295a190a4fa36afcafb241d2ba36c5ce772fd8e0a28eaeeb3e

Observation ef6a9a63-4feb-4fd4-a5d0-f95b4d1712c8 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:56.705569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:56.705569Z digest=sha256:338d0b384cab64487d2333ba2ceb33fdfc8de7a2fce6ed4587813b7f4003e5cf

Observation 5f4e0736-afbd-453c-b5bd-4f278a120af5 · outbound

This paper cites MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:56.774069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:56.774069Z digest=sha256:e99768a9b2a4ff7d73143014a1b2c35cac5468d85070c0491dfecb651268aba5

Observation 624fcccf-52f1-410b-b0ad-ff0eb381e3a4 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs TinyLlama: An Open-Source Small Language Model

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:56.820599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:56.820599Z digest=sha256:bd82a7532967d0f06585e6842d5f42a366e079a4b948028561f2017817a2a71a

Observation aa86406b-5afc-4594-8f49-03318f5ca37f · outbound

This paper cites Group sequential two-stage preference designs.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Group sequential two-stage preference designs

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T20:21:57.960452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:56.870666Z digest=sha256:9f8f0bc731949f991ae9aeb9e76349419f01591cbe2257cf172f735d46915c5b

Observation 4edcb21a-6b65-423f-b4b2-e5c30134afc4 · outbound

This paper cites Pytorch fsdp: Experiences on scaling fully sharded data parallel.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Pytorch fsdp: Experiences on scaling fully sharded data parallel

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:56.984969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:56.984969Z digest=sha256:7fb38116b62deb50f7982fb05b03ca3c4db739226edefadb71d48f3ebc5e48fd

Observation 28521949-bc86-410a-9636-6117c73de4e8 · outbound

This paper cites Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:57.043953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:57.043953Z digest=sha256:347b20ec40d4a8d7b4ef7b1b215d419ad9c52c76681c29d3d286472d798cb784

Observation d743a0a9-bce4-4c02-ad0f-4ad39613fae5 · outbound

This paper cites Toolqa: A dataset for llm question answering with external tools.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Toolqa: A dataset for llm question answering with external tools

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:57.113184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:57.113184Z digest=sha256:6c42f9974972ede14c313b6d4be5a785574be12967ad8212ab01af1a92f7c415

Observation 2954b039-d9ba-498e-9796-f1ae509ed60c · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:57.219298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:57.219298Z digest=sha256:e697b27b35f356e643c0953f895b01bcfdea4a200f8e75b51f5807b473a69f29

Observation 3e516554-032e-4ad8-82c6-58e530a47110 · outbound

This paper cites write newline.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs write newline

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:57.288915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:57.288915Z digest=sha256:f6159d779c692e4f213906bbbb582d5ceffb010d7b6268c7dead0baf5cf2c2b5

Observation eee9aea3-38cd-472f-aa59-e505ecafc354 · outbound

This paper cites @esa (Ref.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs @esa (Ref

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:57.374148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:57.374148Z digest=sha256:c479433f295caf202e5f12a69f76879339ac58d3ee30ef156d05a0753d83a6d2

Observation 71fe0797-da08-430e-a29c-a50b5cbc428a · outbound

This paper cites an unresolved cited work.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:57.452061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:57.452061Z digest=sha256:b3aed3b52ad35021303afddff80654531e0ce401cba87458e6033f9b6239954b

Observation 9887537d-4395-4cc1-8b67-4b122c5a6293 · outbound

This paper cites an unresolved cited work.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:57.501475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:57.501475Z digest=sha256:d344a616008e7159fc8b880c28ab54764b494e4ac2d72477bb1a73ecc44363f8

Pith citing papers

Observation a503399e-fa32-4421-87f0-e1c0935992fe · inbound

IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools cites this paper.

IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:33:58.452689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T05:33:30.670201Z digest=sha256:d289d9e91bd37f16d233403de7578cc9dc9d9b9eca85ff64c088a0034b15c7d5

Observation 4168c4a2-291d-43a3-a6fd-acef3dc18f4c · inbound

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data cites this paper.

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T07:01:44.900291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T07:01:44.900291Z digest=sha256:6dcabbbb434591801b54d7449c6d0b534aa46b3ae7d961d9088c8a672c381270