Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:49:33.125140Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 3 inbound Pith citation observations for arXiv:2505.00358.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:49:33.125140Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:02:00.932716Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T17:58:46.760870Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 81d4a51a-0e1d-4e50-ade2-f2ab54d600d5 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training DoGE: Domain Reweighting with Generalization Estimation
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 224d3fe8-7e53-490e-b635-2bc8bc3f5401 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19297c32-4a8a-4099-8b3b-f3ec652e8c0b · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd02610f-40ee-4562-8007-414a59b4bc28 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Aioli: A Unified Optimization Framework for Language Model Data Mixing
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2458274-30be-4b2f-acce-943e628953b7 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c628917d-b185-45e4-b65a-54a1a1688f7e · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Organize the Web: Constructing Domains Enhances Pre-Training Data Curation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a4d3a7-b7fa-4844-a771-899d6cd21c7d · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5a240039-300b-46bc-8c9f-7b94873a0222 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training DsDm: Model-Aware Dataset Selection with Datamodels
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 432ae583-2d6a-4763-9dc0-3ca9eee05459 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training LESS: Selecting Influential Data for Targeted Instruction Tuning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e330450b-43a4-4a61-81a1-b90d6253c7d4 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Grad-match: Gradient matching based data subset selection for efficient deep model training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5a4f720f-a99b-4093-a136-3438962ed0bf · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f192524-0fa1-4b23-aee7-f9d5116fa2a9 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d5ff941e-fdff-4311-bbc3-e3f5bed18bc8 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Multimodal Data Curation via Object Detection and Filter Ensembles
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c4c4724-04a7-4534-a067-485d6611f390 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Data Selection for Language Models via Importance Resampling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9ede20b-ab34-4238-8082-4b9e67acb531 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training SemDeDup: Data-efficient learning at web-scale through semantic deduplication
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fdad599-8cb4-4361-8f48-a8eb11216028 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Deduplicating Training Data Makes Language Models Better
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 328c1cbe-0204-4b6e-9cc3-b01538cb663f · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training D4: Improving LLM Pretraining via Document De-Duplication and Diversification
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02797f87-ccbb-419c-9662-701610fd8a6b · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training BiMix: A Bivariate Data Mixing Law for Language Model Pretraining
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba23610f-decb-47cb-a908-b3ffb0178216 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 389e68ba-317f-446b-935c-19e3f599e2a2 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training RegMix: Data Mixture as Regression for Language Model Pre-training
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d28db4c1-d9fd-4a04-b58c-501cb6461167 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training AutoScale: Automatic Prediction of Compute-optimal Data Composition for Training LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efaf0881-acb5-401e-9f28-c0bb87518940 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Compute Optimal Scaling of Skills: Knowledge vs Reasoning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 021afc3b-cc65-4b7f-9645-ae37893ef349 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Nomic Embed: Training a Reproducible Long Context Text Embedder
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7099e4e5-4518-412d-bd37-d7c940c624d0 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57af264c-041f-4fc4-b793-ee6f426326b4 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training s1: Simple test-time scaling
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70ca7e16-73e5-408b-9340-8967cb65444f · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 17f27a45-e219-48db-979c-f4f7ed806464 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Dynamic Gradient Alignment for Online Data Mixing
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 789e55bd-e8ab-4a4a-ab52-77deb021e3ce · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38521ce1-5d5a-4fd5-88ac-25b1b63aca47 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Qwen2 Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e16b524e-817a-450d-a245-3435cf3fb005 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; others Learning transferable visual models from natural language supervision
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aefc898a-7480-4144-bb5b-78c81c0fe048 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training OpenCLIP
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb9ef74-0e31-4665-aa46-c9fb67b99a2a · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation db1c332a-6950-4207-b627-174cbfffcb5f · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Scaling Laws for Neural Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14ba2936-c0c3-4606-bf87-a699f693d867 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training What’s the Backward-Forward FLOP Ratio for Neural Networks? 2021; https: //epoch.ai/blog/backward-forward-FLOP-ratio
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9972a7e5-bacd-437a-9d6b-82d0697ed2a6 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Efficient Per-Example Gradient Computations
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a8dc2cd-c182-4726-a852-b732df22b334 · outbound
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training T.; Wu, T.; Song, D.; Mittal, P.; Jia, R
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation df0770f5-85c7-420a-b891-2cf29453ced8 · inbound
Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aab1ca16-6c06-452b-b232-3451f46e26be · inbound
Data Mixing for Large Language Models Pretraining: A Survey and Outlook R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b81392e0-504d-4e37-90eb-d8aa1db6d013 · inbound
WARP: Weight-Space Analysis for Recovering Training Data Portfolios R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.