Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:21:57.501475Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 2 inbound Pith citation observations for arXiv:2507.03253.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:21:57.501475Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T07:01:44.900291Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T05:33:58.451084Z
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7450d016-0969-4188-afac-599e12bcdb68 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79574b7c-16fa-4f40-8ca6-671d9b6310ec · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8fa02c83-c199-41cc-879a-1af021779bf1 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs The claude 3 model family: Opus, sonnet, haiku
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d44f2571-9608-44ad-848e-a9abfb514dac · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Program Synthesis with Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a620327-9de0-4ad6-8008-923297ebb377 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs A critical analysis of the largest source for generative ai training data: Common crawl
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53b6e27e-e9fe-4b74-b5df-c96e9cf58780 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90a0aa22-8055-4f12-8749-da15131acabe · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Piqa: Reasoning about physical commonsense in natural language
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f531adb-7629-4bee-b65c-538d2d47463b · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs On the resemblance and containment of documents
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f50b8dbe-fa86-489f-8281-d08ba7434a6b · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Evaluating Large Language Models Trained on Code
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c03c560-6257-464d-a824-1f44df122b5a · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f456238-d246-4437-b4f1-387d691d94ce · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Flash A ttention-2: Faster attention with better parallelism and work partitioning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca42f1cc-ecfb-413a-946a-75339c94f4d8 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Sailor: Open Language Models for South-East Asia
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00a45189-c2c1-48c3-98c3-a66350e3491d · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b117b13d-bb9a-4808-8711-8260ab07d018 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Minedojo: Building open-ended embodied agents with internet-scale knowledge
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7dc05f53-411e-40c0-bacc-be4be248d23d · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Lighteval: A lightweight framework for llm evaluation, 2023
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 940cbd60-b228-40c5-aa52-27d22bb6a829 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs InCoder: A Generative Model for Code Infilling and Synthesis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0b5aa86-d09c-4d75-bbd7-699cf929be26 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Measuring mathematical problem solving with the math dataset
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b30ca208-ed28-44c2-82af-2ae001040afb · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions, 2023
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc1bd634-0ab6-49a8-9a90-94e64049b0bb · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Gonzalez, Hao Zhang, and Ion Stoica
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be8874e2-b0a5-42da-9f9e-cffd68cd07d8 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Rho-1: Not All Tokens Are What You Need
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 165cd0f0-ee77-4662-8d89-b2ee17be86e4 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Efficient Inference for Large Reasoning Models: A Survey
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 176c9530-0d86-42cd-81b3-64832d10da35 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs OpenELM: An Efficient Language Model Family with Open Training and Inference Framework
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bebc2c8-8413-4360-9d74-d46b248dce4f · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs SLANG: New Concept Comprehension of Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6c9158a-dbfe-4a3a-9b99-50336ebd57b9 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Introducing meta llama 3: The most capable openly available llm to date, 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ae6e40a-e297-4276-b15c-456dbc3e33c2 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Can a suit of armor conduct electricity? a new dataset for open book question answering
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa198cee-3b96-4822-b4c5-75c112b9d7d1 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b978836-5c46-4da4-a3f6-08f365f46dad · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Generative agents: Interactive simulacra of human behavior
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be0ad10-5659-45d7-85d6-078a19e0f096 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b214c50-b851-4640-a7d8-d2f1d12bb48b · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Datatrove: large scale data processing, 2024 b
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef8dcdae-7b02-4ae4-b247-3dc043503be2 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2972caf-9d02-45b9-b472-b5c476447db3 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs DataMan: Data Manager for Pre-training Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b88ea0e-48fe-48ed-a7a7-911089c81335 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df44e4de-4cbc-40f9-944c-b146bd7f49f8 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9babaecc-3f96-47d9-8570-8ce6e469ea9c · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Web data mining with organized contents using naive bayes algorithm
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4fea4167-8318-4f55-8c65-2fc992a212fc · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Winogrande: An adversarial winograd schema challenge at scale
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de97d225-d551-42fc-9daf-9ec83628ca63 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs SocialIQA: Commonsense Reasoning about Social Interactions
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75c4317c-ccea-4174-a1ff-5c9aa60e5590 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Toolformer: Language models can teach themselves to use tools
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9621e435-1282-4009-87e8-92e5d1623dc4 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e0e0e2c-b9bb-4f69-897c-f218fddc8c6f · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs SlimPajama: A 627B token cleaned and deduplicated version of RedPajama , June 2023
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f9a3b0d5-5a17-4bd4-a6da-0334819f2b0c · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Dolma: an open corpus of three trillion tokens for language model pretraining research
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 34c86beb-14da-4524-bff9-e9b876b99569 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e3606bc-1499-4c96-8dd5-121609bd9e60 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs C ommonsense QA : A question answering challenge targeting commonsense knowledge
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b1017bf-c110-4a24-9483-2bccce464167 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Redpajama: an open dataset for training large language models, October 2023
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0e9793d1-3fec-4858-a7bf-134458de5c9b · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs M Towhidul Islam Tonmoy, S M Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8738f8af-7e15-4c73-8233-0a045088789e · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs LLaMA: Open and Efficient Foundation Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d90a520-0610-4b87-af93-cbc33ab5fd3e · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c47027db-4378-4261-9fd5-3943afbc1431 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Code Llama: Open Foundation Models for Code
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41e22c1e-de88-4907-9578-997b904eb7bf · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Chain-of-thought prompting elicits reasoning in large language models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f417b6c-ba9a-4484-8601-42c4466c8828 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Crowdsourcing Multiple Choice Science Questions
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 594b4f42-1bf6-46a6-9157-bd1f6a7078d3 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs QuRating : Selecting high-quality data for training language models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c19a7082-0c5e-4986-893e-1ed92657c4fb · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Data selection for language models via importance resampling
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 692d263e-4ffd-4a57-a9c3-087be751aa98 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Qwen3 Technical Report
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d2a5181-c345-4c75-a198-2b59d2c5aafd · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs React: Synergizing reasoning and acting in language models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 591e6031-5101-4b14-9a15-3885759666b3 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Craw4LLM: Efficient Web Crawling for LLM Pretraining
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3e0ccfe-6668-4d09-b5f5-7a27ce0e6b6c · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef9f0231-9922-4d4c-b621-f86de036139a · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs A normalized levenshtein distance metric
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef6a9a63-4feb-4fd4-a5d0-f95b4d1712c8 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f4e0736-afbd-453c-b5bd-4f278a120af5 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 624fcccf-52f1-410b-b0ad-ff0eb381e3a4 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs TinyLlama: An Open-Source Small Language Model
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa86406b-5afc-4594-8f49-03318f5ca37f · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Group sequential two-stage preference designs
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4edcb21a-6b65-423f-b4b2-e5c30134afc4 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Pytorch fsdp: Experiences on scaling fully sharded data parallel
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28521949-bc86-410a-9636-6117c73de4e8 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d743a0a9-bce4-4c02-ad0f-4ad39613fae5 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Toolqa: A dataset for llm question answering with external tools
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2954b039-d9ba-498e-9796-f1ae509ed60c · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e516554-032e-4ad8-82c6-58e530a47110 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs write newline
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eee9aea3-38cd-472f-aa59-e505ecafc354 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs @esa (Ref
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71fe0797-da08-430e-a29c-a50b5cbc428a · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Unresolved cited work
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9887537d-4395-4cc1-8b67-4b122c5a6293 · outbound
RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs Unresolved cited work
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a503399e-fa32-4421-87f0-e1c0935992fe · inbound
IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4168c4a2-291d-43a3-a6fd-acef3dc18f4c · inbound
DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.