Pith. sign in

Paper Citation Record · LEDGER

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs

As of 7 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 2 inbound Pith citation observations for arXiv:2507.03253.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.03253 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:21:57.501475Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T07:01:44.900291Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T05:33:58.451084Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact3
  • verified fuzzy17
  • unresolved46
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7450d016-0969-4188-afac-599e12bcdb68 · outbound

This paper cites GPT-4 Technical Report.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:49.221314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:49.221314Z digest=sha256:9b90ce965a45b8f265ab05c7555e186894629a65634dcf166d775d9bdcc8199a

Observation 79574b7c-16fa-4f40-8ca6-671d9b6310ec · outbound

This paper cites an unresolved cited work.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:22:01.871394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:49.280948Z digest=sha256:156abd2517de0d5871085a2edb8b102a73efb4856a0ab5587ab671d7ae1c10e4

Observation 8fa02c83-c199-41cc-879a-1af021779bf1 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs The claude 3 model family: Opus, sonnet, haiku

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:01.684652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:49.382321Z digest=sha256:4aca022168d217ec70733aefe38c0f7ce404917560e06f96993c3b69c6859b92

Observation d44f2571-9608-44ad-848e-a9abfb514dac · outbound

This paper cites Program Synthesis with Large Language Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Program Synthesis with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:49.508104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:49.508104Z digest=sha256:0264d00930d94fc5e9ad5dce899a592acd6b3cd2bf0a004b8d57f5fcae90082b

Observation 2a620327-9de0-4ad6-8008-923297ebb377 · outbound

This paper cites A critical analysis of the largest source for generative ai training data: Common crawl.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs A critical analysis of the largest source for generative ai training data: Common crawl

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:01.574642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:49.662920Z digest=sha256:c20e185b79cb6cc9b1e3d8754aebc867f40fb3ae34a75cf75c03b16dde0da09d

Observation 53b6e27e-e9fe-4b74-b5df-c96e9cf58780 · outbound

This paper cites Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:49.842658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:49.842658Z digest=sha256:d5e0bc1e28d4dd3a7cd697a88deb8250596fbcc2a15233e550b728c111709692

Observation 90a0aa22-8055-4f12-8749-da15131acabe · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Piqa: Reasoning about physical commonsense in natural language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:49.980601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:49.980601Z digest=sha256:a314b0d24723446a0eee32830145273f13a7c59b148e5faa9aea1661ed4daa11

Observation 7f531adb-7629-4bee-b65c-538d2d47463b · outbound

This paper cites On the resemblance and containment of documents.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs On the resemblance and containment of documents

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:01.424074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:50.161331Z digest=sha256:da5d3bf44fec0f1ed3f1f5cddd08975dea86fa44bcfc4102740dd9fc93020300

Observation f50b8dbe-fa86-489f-8281-d08ba7434a6b · outbound

This paper cites Evaluating Large Language Models Trained on Code.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Evaluating Large Language Models Trained on Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:50.343916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:50.343916Z digest=sha256:2b093f0f6c69b7ec2208f5e19fd90444be37fe2c3b4489892f3aab9bc2319e5a

Observation 8c03c560-6257-464d-a824-1f44df122b5a · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:50.502340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:50.502340Z digest=sha256:015818183f27c69b3f55d2b7df114db14290f8d22bff84dcb1f8b003006fe936

Observation 1f456238-d246-4437-b4f1-387d691d94ce · outbound

This paper cites Flash A ttention-2: Faster attention with better parallelism and work partitioning.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Flash A ttention-2: Faster attention with better parallelism and work partitioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:50.621007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:50.621007Z digest=sha256:79d63c7a364ff159415eadfb7e9bcf03fb6b70613a9b53d044ca64b4f584c929

Observation ca42f1cc-ecfb-413a-946a-75339c94f4d8 · outbound

This paper cites Sailor: Open Language Models for South-East Asia.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Sailor: Open Language Models for South-East Asia

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:21:57.715855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:50.740610Z digest=sha256:83c5ff816920c742d51954fb15b0991f0e73141d092a753102a7144754fb1b18

Observation 00a45189-c2c1-48c3-98c3-a66350e3491d · outbound

This paper cites The Llama 3 Herd of Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:50.905266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:50.905266Z digest=sha256:899d0aea23289aaf5b3ea89aed9752bbf8c82b7c50c361a3f744eac3e615b22f

Observation b117b13d-bb9a-4808-8711-8260ab07d018 · outbound

This paper cites Minedojo: Building open-ended embodied agents with internet-scale knowledge.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Minedojo: Building open-ended embodied agents with internet-scale knowledge

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:01.267272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:51.070802Z digest=sha256:e258b182ab553e0786de93ec248921b32dd4901908ab8f5885e6ff097a0f1968

Observation 7dc05f53-411e-40c0-bacc-be4be248d23d · outbound

This paper cites Lighteval: A lightweight framework for llm evaluation, 2023.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Lighteval: A lightweight framework for llm evaluation, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:01.120985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:51.224440Z digest=sha256:9628e20fc3431849e70af9db6a0a6ec567885cd18f1d6bd37a25f9345ae9f08f

Observation 940cbd60-b228-40c5-aa52-27d22bb6a829 · outbound

This paper cites InCoder: A Generative Model for Code Infilling and Synthesis.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs InCoder: A Generative Model for Code Infilling and Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:51.342397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:51.342397Z digest=sha256:6a79c6c2a988e14a71cac7a1e680e5fd45a77b15133e407f94bf8baededf3bb3

Observation d0b5aa86-d09c-4d75-bbd7-699cf929be26 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Measuring mathematical problem solving with the math dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:51.486728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:51.486728Z digest=sha256:060ca7c2db05862a940fbe5d922a95c466b1f32716d3388be7528d27d43c229b

Observation b30ca208-ed28-44c2-82af-2ae001040afb · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions, 2023.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:00.978871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:51.629900Z digest=sha256:9f2a69084907840fb658c8bd438ff5742bec04d581b07029212ea5acab971dee

Observation fc1bd634-0ab6-49a8-9a90-94e64049b0bb · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Gonzalez, Hao Zhang, and Ion Stoica

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:51.719872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:51.719872Z digest=sha256:16f67282d7782fec69f0c6f56a5d771f6be99d75672a48c3a43d6a4913884914

Observation be8874e2-b0a5-42da-9f9e-cffd68cd07d8 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Rho-1: Not All Tokens Are What You Need

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:51.821393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:51.821393Z digest=sha256:337c49d5ac909ae1c2c32f46465e13779cab1b771001957cc722889538d6a582

Observation 165cd0f0-ee77-4662-8d89-b2ee17be86e4 · outbound

This paper cites Efficient Inference for Large Reasoning Models: A Survey.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Efficient Inference for Large Reasoning Models: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:51.911641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:51.911641Z digest=sha256:580031c64a5f5c0eb493044b48a9aa536a2da51fb7b8a8c1c7a880fe3726859a

Observation 176c9530-0d86-42cd-81b3-64832d10da35 · outbound

This paper cites OpenELM: An Efficient Language Model Family with Open Training and Inference Framework.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs OpenELM: An Efficient Language Model Family with Open Training and Inference Framework

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:51.990657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:51.990657Z digest=sha256:efaa0943ad8e47477aa4507cdd4c0ecb7432774945b8fe72daaf9e9d3dc2c6e4

Observation 8bebc2c8-8413-4360-9d74-d46b248dce4f · outbound

This paper cites SLANG: New Concept Comprehension of Large Language Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs SLANG: New Concept Comprehension of Large Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:21:58.461364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:52.082429Z digest=sha256:ea4d9cadabd42c1ab8e76991419b8c62718dab6d2e48f1b400f949ebfe95ad9c

Observation f6c9158a-dbfe-4a3a-9b99-50336ebd57b9 · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date, 2024.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Introducing meta llama 3: The most capable openly available llm to date, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:00.766841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:52.176927Z digest=sha256:50de537e529f40f20ee406951b2694bc023260f2199549dcd9aefd232ad25aff

Observation 9ae6e40a-e297-4276-b15c-456dbc3e33c2 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:52.245731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:52.245731Z digest=sha256:884a4e4d682d44abc6a3eed765858341d9cec3d81d933d4ad372a7f484d3168f

Observation aa198cee-3b96-4822-b4c5-75c112b9d7d1 · outbound

This paper cites Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:52.353944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:52.353944Z digest=sha256:5fb11e3ea0f875c470e628f49e0f30fe708d46a9455687051f54c996506d7fe0

Observation 5b978836-5c46-4da4-a3f6-08f365f46dad · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Generative agents: Interactive simulacra of human behavior

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:52.474419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:52.474419Z digest=sha256:38fa2c695149f1405d58cf16428a932d6543bab27629b752f7be03d6c3b9b0b2

Observation 3be0ad10-5659-45d7-85d6-078a19e0f096 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:52.628204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:52.628204Z digest=sha256:44b244862544eff604198f2ccb8826e4316b4a8ac65f38f5560901162ad6b9e5

Observation 3b214c50-b851-4640-a7d8-d2f1d12bb48b · outbound

This paper cites Datatrove: large scale data processing, 2024 b.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Datatrove: large scale data processing, 2024 b

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:00.595823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:52.744819Z digest=sha256:c2ca02f16c082b703039c304b0d8a5914ac152e6453e8009dfc0f97c7020553c

Observation ef8dcdae-7b02-4ae4-b247-3dc043503be2 · outbound

This paper cites The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:00.412211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:52.879117Z digest=sha256:06a2099ee16dbdb9f24cd1250b1dbd3641490e1b332d79005d24d1425fa05915

Observation e2972caf-9d02-45b9-b472-b5c476447db3 · outbound

This paper cites DataMan: Data Manager for Pre-training Large Language Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs DataMan: Data Manager for Pre-training Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:53.021981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:53.021981Z digest=sha256:99967d2ec4916c905d21dbee6ff64cf3e0d671cb43b508ca82a8e3aa3f5b19dd

Observation 2b88ea0e-48fe-48ed-a7a7-911089c81335 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:53.148940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:53.148940Z digest=sha256:c2206a6914e7dc6d93aa9dbed775f7b81b5e6154308253e5b9a0ae9f5d4534ca

Observation df44e4de-4cbc-40f9-944c-b146bd7f49f8 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:53.276242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:53.276242Z digest=sha256:41d53e3b3dc2f241dc3343674137730cc75b7b28dd2ed52d6fb50fde0d63b01c

Observation 9babaecc-3f96-47d9-8570-8ce6e469ea9c · outbound

This paper cites Web data mining with organized contents using naive bayes algorithm.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Web data mining with organized contents using naive bayes algorithm

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:00.139969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:53.406768Z digest=sha256:7eb11cfa9f93b0c7f8e5e21c2e50f361927231a61980edaa903f0c604ff2ca93

Observation 4fea4167-8318-4f55-8c65-2fc992a212fc · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Winogrande: An adversarial winograd schema challenge at scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:53.567288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:53.567288Z digest=sha256:68e883f16760c07e79de3e310bc514f98e6ebe50be092000c6f43ba545ebbcd2

Observation de97d225-d551-42fc-9daf-9ec83628ca63 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs SocialIQA: Commonsense Reasoning about Social Interactions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:53.691167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:53.691167Z digest=sha256:99b501a9c4801cb6fdfcfe1ad125b137736b9ec837d61ae5ce47791fa7f3d7e4

Observation 75c4317c-ccea-4174-a1ff-5c9aa60e5590 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Toolformer: Language models can teach themselves to use tools

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:22:00.002087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:53.844678Z digest=sha256:276bf00acf151cb0927420ebcd68a0aee3023b70b78e29111ce6f7b053e65105

Observation 9621e435-1282-4009-87e8-92e5d1623dc4 · outbound

This paper cites Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:53.998507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:53.998507Z digest=sha256:feae5f3a1583b0a82deb12ca0430c3a85f2fb46597d5eb54b0293164a5cc167f

Observation 4e0e0e2c-b9bb-4f69-897c-f218fddc8c6f · outbound

This paper cites SlimPajama: A 627B token cleaned and deduplicated version of RedPajama , June 2023.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs SlimPajama: A 627B token cleaned and deduplicated version of RedPajama , June 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:21:59.833764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:54.173260Z digest=sha256:63be418bf7bdeb9e30a357998c0ce95f2dc245c11129c717e2ebe846983ca241

Observation f9a3b0d5-5a17-4bd4-a6da-0334819f2b0c · outbound

This paper cites Dolma: an open corpus of three trillion tokens for language model pretraining research.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Dolma: an open corpus of three trillion tokens for language model pretraining research

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:21:59.545747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:54.325158Z digest=sha256:a146bafd2ba3f7d45e31ac1ed64968f9907023f43ee9f1b75119e01286b41fa7

Observation 34c86beb-14da-4524-bff9-e9b876b99569 · outbound

This paper cites Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:54.439779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:54.439779Z digest=sha256:9f31353bee4246bfa0f886f703b41c9a59b9c8bd33a74daacc09c95e6082ec6a

Observation 3e3606bc-1499-4c96-8dd5-121609bd9e60 · outbound

This paper cites C ommonsense QA : A question answering challenge targeting commonsense knowledge.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs C ommonsense QA : A question answering challenge targeting commonsense knowledge

Reference 42

Resolution
malformed identifier
no resolver link, observed 2026-08-06T20:21:54.624141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:54.624141Z digest=sha256:effe2a5392158bd13cae11e7585c7fd28b6bee751613a2aa82b9482dfe8a28d9

Observation 5b1017bf-c110-4a24-9483-2bccce464167 · outbound

This paper cites Redpajama: an open dataset for training large language models, October 2023.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Redpajama: an open dataset for training large language models, October 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:21:59.401589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:54.781545Z digest=sha256:03b7cca08f072b0655b6f0a7f1bdcb663bf8c830d97491290663d7a7bf6e4939

Observation 0e9793d1-3fec-4858-a7bf-134458de5c9b · outbound

This paper cites M Towhidul Islam Tonmoy, S M Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs M Towhidul Islam Tonmoy, S M Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:21:59.219332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:54.951767Z digest=sha256:e1175bee94610ce572ac232e26ac788d160b5c0091e3338980fb1ddfe4fa5a01

Observation 8738f8af-7e15-4c73-8233-0a045088789e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs LLaMA: Open and Efficient Foundation Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:55.088543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:55.088543Z digest=sha256:8acc16b101c47a1cf24225d7b39a6ff020d996c9b2bb3ec525be3b5b2e76a6a2

Observation 5d90a520-0610-4b87-af93-cbc33ab5fd3e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:55.236194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:55.236194Z digest=sha256:d9c8caf554a4a92ddcd9a80f538ba5634ebb04d35bf000f41a96c583aeeb3f60

Observation c47027db-4378-4261-9fd5-3943afbc1431 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Code Llama: Open Foundation Models for Code

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:55.411568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:55.411568Z digest=sha256:4c8ac0b487c85eebefc795ff8169659f825b36bc4ca2226982aa42b49ad416be

Observation 41e22c1e-de88-4907-9578-997b904eb7bf · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Chain-of-thought prompting elicits reasoning in large language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:55.555671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:55.555671Z digest=sha256:f0b4cae740eb1e0374ab6cd4ab34c2ed8d1befc2ad821320e512b6514647e054

Observation 7f417b6c-ba9a-4484-8601-42c4466c8828 · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Crowdsourcing Multiple Choice Science Questions

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:55.665550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:55.665550Z digest=sha256:ed335ae6ae114ce9b3378d85c2c492785c2496bc40311d878f71d9ec3697f607

Observation 594b4f42-1bf6-46a6-9157-bd1f6a7078d3 · outbound

This paper cites QuRating : Selecting high-quality data for training language models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs QuRating : Selecting high-quality data for training language models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:55.821166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:55.821166Z digest=sha256:8eac0a9b04d58a8d498de22bbf68239608358af0da589995ed6d091b73732de2

Observation c19a7082-0c5e-4986-893e-1ed92657c4fb · outbound

This paper cites Data selection for language models via importance resampling.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Data selection for language models via importance resampling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:55.957694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:55.957694Z digest=sha256:b8757966030f23613b70f547a513577ef8d1050dca31cf01e95ab654d4470234

Observation 692d263e-4ffd-4a57-a9c3-087be751aa98 · outbound

This paper cites Qwen3 Technical Report.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Qwen3 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:56.118676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:56.118676Z digest=sha256:1122efa3fa31bce199f21b34b7b9e66738ac945ea45d73b2b617dab81016ad26

Observation 8d2a5181-c345-4c75-a198-2b59d2c5aafd · outbound

This paper cites React: Synergizing reasoning and acting in language models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs React: Synergizing reasoning and acting in language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:21:58.985198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:56.290379Z digest=sha256:33c94e2eff2bbe5cf334bd1f2b936861a5a321f563b442acd5a19ae82fae90d5

Observation 591e6031-5101-4b14-9a15-3885759666b3 · outbound

This paper cites Craw4LLM: Efficient Web Crawling for LLM Pretraining.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Craw4LLM: Efficient Web Crawling for LLM Pretraining

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:21:58.169963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:56.432140Z digest=sha256:eacd6d1bb081ef697161a98131f58928c64e2eacb3843f45ef3e9b1efc61b39a

Observation f3e0ccfe-6668-4d09-b5f5-7a27ce0e6b6c · outbound

This paper cites MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:56.521904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:56.521904Z digest=sha256:731c4e1c6ace7728cb3cae35a47617ec335d7251ca24dbd0a4dba5d43e8760a7

Observation ef9f0231-9922-4d4c-b621-f86de036139a · outbound

This paper cites A normalized levenshtein distance metric.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs A normalized levenshtein distance metric

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:21:58.795332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:56.621387Z digest=sha256:f54147eeb37caf649f6815bee74891363350a17c13a38498d0b0d4f985fb3bf3

Observation ef6a9a63-4feb-4fd4-a5d0-f95b4d1712c8 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:56.705569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:56.705569Z digest=sha256:0993d9a82641666c9cb6c626da457b7fb92daf7a35d4ff62cf05000354c7ed40

Observation 5f4e0736-afbd-453c-b5bd-4f278a120af5 · outbound

This paper cites MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:56.774069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:56.774069Z digest=sha256:00716db8e5d544e456264749574c24a13e165f8372e8042640c45fa479d98cb9

Observation 624fcccf-52f1-410b-b0ad-ff0eb381e3a4 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs TinyLlama: An Open-Source Small Language Model

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:56.820599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:56.820599Z digest=sha256:41a8d1150c8189c9b7bd980d3c10c4faa94088e5ae313129c4d6939cef9f85f0

Observation aa86406b-5afc-4594-8f49-03318f5ca37f · outbound

This paper cites Group sequential two-stage preference designs.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Group sequential two-stage preference designs

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T20:21:57.960452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T20:21:56.870666Z digest=sha256:80e0a8250b4be8a46cd27a5c361fce96b440ed5c0db1e733771f2e1eaf4cb58f

Observation 4edcb21a-6b65-423f-b4b2-e5c30134afc4 · outbound

This paper cites Pytorch fsdp: Experiences on scaling fully sharded data parallel.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Pytorch fsdp: Experiences on scaling fully sharded data parallel

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:56.984969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:56.984969Z digest=sha256:4c3329d7a4eb842f080d37e55bf35a0596898485a253d8d5f0dadde7863edce8

Observation 28521949-bc86-410a-9636-6117c73de4e8 · outbound

This paper cites Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:57.043953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:57.043953Z digest=sha256:4692d3bf5bf8a43b8e46d77cfb5b19766f959b553cbbadb9d2780a9d145d11fa

Observation d743a0a9-bce4-4c02-ad0f-4ad39613fae5 · outbound

This paper cites Toolqa: A dataset for llm question answering with external tools.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Toolqa: A dataset for llm question answering with external tools

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:57.113184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:57.113184Z digest=sha256:2c30cac58114b64c3184b8e09be34f65a84eeb549ef58d1b0a0ac46416d18665

Observation 2954b039-d9ba-498e-9796-f1ae509ed60c · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:57.219298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:57.219298Z digest=sha256:0fa0eead00a030412f81014f8134838d3cac61ab4c7dce8041cd06ac77e33ff0

Observation 3e516554-032e-4ad8-82c6-58e530a47110 · outbound

This paper cites write newline.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs write newline

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:57.288915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:57.288915Z digest=sha256:8b56e902f52213b46081dc1c29b392740685929acd3184a6afd67865f8c991ce

Observation eee9aea3-38cd-472f-aa59-e505ecafc354 · outbound

This paper cites @esa (Ref.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs @esa (Ref

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:57.374148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:57.374148Z digest=sha256:947bfdf9d3d8aaf91a5b0a97d87c63d76a4c496578de570d12f9b96a70d1852e

Observation 71fe0797-da08-430e-a29c-a50b5cbc428a · outbound

This paper cites an unresolved cited work.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:57.452061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:57.452061Z digest=sha256:88a37b53da0b7cd3d07c6cf2c62cfcdc4681e3215292c5d6d64bdc0c3966b451

Observation 9887537d-4395-4cc1-8b67-4b122c5a6293 · outbound

This paper cites an unresolved cited work.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:57.501475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:57.501475Z digest=sha256:0344e76bcf36cb6d69b0e05ed35223e116b2ab15c243e4303502105cbb0db6a2

Pith citing papers

Observation a503399e-fa32-4421-87f0-e1c0935992fe · inbound

IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools cites this paper.

IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:33:58.452689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T05:33:30.670201Z digest=sha256:d2032d71fd8278b8f26b8b70415296f14a6fad013fa4a06f3ac913703829e0f6

Observation 4168c4a2-291d-43a3-a6fd-acef3dc18f4c · inbound

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data cites this paper.

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T07:01:44.900291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T07:01:44.900291Z digest=sha256:a3eb46825957accbb95c22b9fcf975bcdccce9e58c6e69dfd288f5cc412bcdef