Pith. sign in

Paper Citation Record · LEDGER

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training

As of 22 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 3 inbound Pith citation observations for arXiv:2505.00358.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00358 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:49:33.125140Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:02:00.932716Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:58:46.760870Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 81d4a51a-0e1d-4e50-ade2-f2ab54d600d5 · outbound

This paper cites DoGE: Domain Reweighting with Generalization Estimation.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training DoGE: Domain Reweighting with Generalization Estimation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:32.939849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:32.939849Z digest=sha256:9e02da05b933f3fe8ef09b766b933deb699ac80f2be1c656ff93a45038256630

Observation 224d3fe8-7e53-490e-b635-2bc8bc3f5401 · outbound

This paper cites DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:32.945685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:32.945685Z digest=sha256:22b024d5cb8caac8d91b032134c00f7b582323d5507b883e8eba72e6bd650198

Observation 19297c32-4a8a-4099-8b3b-f3ec652e8c0b · outbound

This paper cites Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:32.951293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:32.951293Z digest=sha256:18b1244d74c4c9de69141b36d8f00a1c3b2b3b59d4554bae97e77253abb7e7ee

Observation dd02610f-40ee-4562-8007-414a59b4bc28 · outbound

This paper cites Aioli: A Unified Optimization Framework for Language Model Data Mixing.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Aioli: A Unified Optimization Framework for Language Model Data Mixing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:32.957200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:32.957200Z digest=sha256:cebbc61ba219ba56fa08c0963df200642861545eb9b7e8cef0cdf6ad92247904

Observation d2458274-30be-4b2f-acce-943e628953b7 · outbound

This paper cites Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:32.962436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:32.962436Z digest=sha256:530f9894327ed99f4a7fa0f16ca696e65bc7590b243bdf7b6439740b6402f6bb

Observation c628917d-b185-45e4-b65a-54a1a1688f7e · outbound

This paper cites Organize the Web: Constructing Domains Enhances Pre-Training Data Curation.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Organize the Web: Constructing Domains Enhances Pre-Training Data Curation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:32.967892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:32.967892Z digest=sha256:0233ffb90b21388b646b0c8672780309913b1cef0ed3b4bfa512b4a9501f8538

Observation 41a4d3a7-b7fa-4844-a771-899d6cd21c7d · outbound

This paper cites Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:49:33.862331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:49:32.973525Z digest=sha256:337c6fe4a1dbc2e906218cbdcda8b7bee0fef7c11f9eb67db131f729bfc0420b

Observation 5a240039-300b-46bc-8c9f-7b94873a0222 · outbound

This paper cites DsDm: Model-Aware Dataset Selection with Datamodels.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training DsDm: Model-Aware Dataset Selection with Datamodels

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:32.978094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:32.978094Z digest=sha256:3655ef28ab206733e3ec1fa54f9d7ce4e5c24bf0e97ea208c1a823d7f2261dd7

Observation 432ae583-2d6a-4763-9dc0-3ca9eee05459 · outbound

This paper cites LESS: Selecting Influential Data for Targeted Instruction Tuning.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training LESS: Selecting Influential Data for Targeted Instruction Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:32.984118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:32.984118Z digest=sha256:e13afdc8d94f7ddbc7535370c7503c760a7059c61fc96a7f5fdeacb7a15d2c72

Observation e330450b-43a4-4a61-81a1-b90d6253c7d4 · outbound

This paper cites Grad-match: Gradient matching based data subset selection for efficient deep model training.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Grad-match: Gradient matching based data subset selection for efficient deep model training

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:49:33.845746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:49:32.989113Z digest=sha256:235f49f0a88c8d1334c61db8c1505735e1847468c1eaf4baf86e3a3a938e51ab

Observation 5a4f720f-a99b-4093-a136-3438962ed0bf · outbound

This paper cites Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:32.994117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:32.994117Z digest=sha256:b3437da884d38adeb92ad22e9d49e4f4842bc1f0a06b63fd0bc3864f42e183af

Observation 9f192524-0fa1-4b23-aee7-f9d5116fa2a9 · outbound

This paper cites Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:49:33.565611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:49:32.999255Z digest=sha256:665dc42ed76e1b32d001ede1f3a5126f1fc2ee7e516ca11106e70e29285e91f5

Observation d5ff941e-fdff-4311-bbc3-e3f5bed18bc8 · outbound

This paper cites Multimodal Data Curation via Object Detection and Filter Ensembles.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Multimodal Data Curation via Object Detection and Filter Ensembles

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.005055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.005055Z digest=sha256:cb375d7a1112605eb8b76a6bebed150354f7d3507e3d76011e2defe41e000eb1

Observation 2c4c4724-04a7-4534-a067-485d6611f390 · outbound

This paper cites Data Selection for Language Models via Importance Resampling.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Data Selection for Language Models via Importance Resampling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.010284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.010284Z digest=sha256:c101dd38dc7be1853706dd7d91c207b1514dcf47209b31c4be1fe206ae5e0718

Observation f9ede20b-ab34-4238-8082-4b9e67acb531 · outbound

This paper cites SemDeDup: Data-efficient learning at web-scale through semantic deduplication.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training SemDeDup: Data-efficient learning at web-scale through semantic deduplication

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.015503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.015503Z digest=sha256:77c57141624a14dff829d113a2f72f065b6fa14dee2c28c3573108a386070cba

Observation 2fdad599-8cb4-4361-8f48-a8eb11216028 · outbound

This paper cites Deduplicating Training Data Makes Language Models Better.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Deduplicating Training Data Makes Language Models Better

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.021069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.021069Z digest=sha256:57bd1ebf475aa26dabaf7262d892815151a39833e0a668e29b5da79b9a63ef03

Observation 328c1cbe-0204-4b6e-9cc3-b01538cb663f · outbound

This paper cites D4: Improving LLM Pretraining via Document De-Duplication and Diversification.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training D4: Improving LLM Pretraining via Document De-Duplication and Diversification

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.026439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.026439Z digest=sha256:fca41b935e2827b5f0b6543e43235d461e0b3b18f467d20420599388503a5849

Observation 02797f87-ccbb-419c-9662-701610fd8a6b · outbound

This paper cites BiMix: A Bivariate Data Mixing Law for Language Model Pretraining.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training BiMix: A Bivariate Data Mixing Law for Language Model Pretraining

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.032610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.032610Z digest=sha256:585b9eb2f364f6f7072533fc4320aaff04958bc06a0ef3ff5eff95b41e25e9ae

Observation ba23610f-decb-47cb-a908-b3ffb0178216 · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.038457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.038457Z digest=sha256:1151f47ec279c901a6f748de6cb536a1486ff4a5722f4284a01cea5ba22004ed

Observation 389e68ba-317f-446b-935c-19e3f599e2a2 · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.044017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.044017Z digest=sha256:e7e96d419ba74b855cbde392a2b44b4d28c3db6364375209e92005452fcbccad

Observation d28db4c1-d9fd-4a04-b58c-501cb6461167 · outbound

This paper cites AutoScale: Automatic Prediction of Compute-optimal Data Composition for Training LLMs.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training AutoScale: Automatic Prediction of Compute-optimal Data Composition for Training LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.048755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.048755Z digest=sha256:9e5f5d6aaf98ba53e8c0929f8b9df7cce40bb6b01d384c01d820c030f134c99a

Observation efaf0881-acb5-401e-9f28-c0bb87518940 · outbound

This paper cites Compute Optimal Scaling of Skills: Knowledge vs Reasoning.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Compute Optimal Scaling of Skills: Knowledge vs Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.053552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.053552Z digest=sha256:23b7c667bed3caf13521830074282902132a5f48920ece57cbe69f25428cd5de

Observation 021afc3b-cc65-4b7f-9645-ae37893ef349 · outbound

This paper cites Nomic Embed: Training a Reproducible Long Context Text Embedder.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Nomic Embed: Training a Reproducible Long Context Text Embedder

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.058764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.058764Z digest=sha256:d68d73a972fb37365724dcd5e6e8fa80e58decb753d86a4d6a5da7a37cb39f0e

Observation 7099e4e5-4518-412d-bd37-d7c940c624d0 · outbound

This paper cites Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.064158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.064158Z digest=sha256:721302a2a63134de1cad17bb0b446a14114b51808fb24664b08485f0ea12fb13

Observation 57af264c-041f-4fc4-b793-ee6f426326b4 · outbound

This paper cites s1: Simple test-time scaling.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training s1: Simple test-time scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.070209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.070209Z digest=sha256:00324f44fd67ec4227c6f63327f5951de1d81f52815eb48261d074b2e8537c3b

Observation 70ca7e16-73e5-408b-9340-8967cb65444f · outbound

This paper cites an unresolved cited work.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:49:33.829305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:49:33.075458Z digest=sha256:793cd775d70a9ae4f2b92a31f4b22136af114165f77dab6a15e1d35c4f768b56

Observation 17f27a45-e219-48db-979c-f4f7ed806464 · outbound

This paper cites Dynamic Gradient Alignment for Online Data Mixing.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Dynamic Gradient Alignment for Online Data Mixing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.080406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.080406Z digest=sha256:2f6bfd4ea831dbbcba3c247777c73dff9980e497772df4aaed8c2341605b0b2d

Observation 789e55bd-e8ab-4a4a-ab52-77deb021e3ce · outbound

This paper cites GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.085127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.085127Z digest=sha256:117329ada7ebaa2cf3a0a8309209593241ea2416b2f95c542598f3b166336499

Observation 38521ce1-5d5a-4fd5-88ac-25b1b63aca47 · outbound

This paper cites Qwen2 Technical Report.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Qwen2 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.090579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.090579Z digest=sha256:d057aab4063566369a1fbaf17df9a7aeb6bf0125fcbafaf728c95af36a7c63cf

Observation e16b524e-817a-450d-a245-3435cf3fb005 · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; others Learning transferable visual models from natural language supervision.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; others Learning transferable visual models from natural language supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.095840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.095840Z digest=sha256:5cc9fc16e8e596d4a5ba846de9b40f3ea877497a7d5d8acb01ce95b92370adfe

Observation aefc898a-7480-4144-bb5b-78c81c0fe048 · outbound

This paper cites OpenCLIP.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training OpenCLIP

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.100308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.100308Z digest=sha256:deed9c6f662a457a02c757f178f6afd36030d4ea79bfa2b2ce294d1e14688fcf

Observation 7fb9ef74-0e31-4665-aa46-c9fb67b99a2a · outbound

This paper cites an unresolved cited work.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:49:33.801543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:49:33.104899Z digest=sha256:5b948f156316da9e2d502ad7b3a3914e7556262fdea2eb521b31b84086b2fa59

Observation db1c332a-6950-4207-b627-174cbfffcb5f · outbound

This paper cites Scaling Laws for Neural Language Models.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Scaling Laws for Neural Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.109757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.109757Z digest=sha256:eaad0fbc4835ced24a3b18359ec7602b3177b6dfc393aa93c5ace00922be34dc

Observation 14ba2936-c0c3-4606-bf87-a699f693d867 · outbound

This paper cites What’s the Backward-Forward FLOP Ratio for Neural Networks? 2021; https: //epoch.ai/blog/backward-forward-FLOP-ratio.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training What’s the Backward-Forward FLOP Ratio for Neural Networks? 2021; https: //epoch.ai/blog/backward-forward-FLOP-ratio

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:49:33.785127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:49:33.114538Z digest=sha256:9457e7b14406e8f36085f27771b63c43b3d8a481b2f6d886d1b170a4cdd522ef

Observation 9972a7e5-bacd-437a-9d6b-82d0697ed2a6 · outbound

This paper cites Efficient Per-Example Gradient Computations.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training Efficient Per-Example Gradient Computations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T04:49:33.119833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:49:33.119833Z digest=sha256:17c733135d1269df3151a2dde83faa94f84f66cde44d5b0244faf7f1cd212286

Observation 1a8dc2cd-c182-4726-a852-b732df22b334 · outbound

This paper cites T.; Wu, T.; Song, D.; Mittal, P.; Jia, R.

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training T.; Wu, T.; Song, D.; Mittal, P.; Jia, R

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:49:33.767439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T04:49:33.125140Z digest=sha256:c78b85a4bc4ed64aab5374ed2e46adc4774017e4230f7b083186f0a72db7b3d8

Pith citing papers

Observation df0770f5-85c7-420a-b891-2cf29453ced8 · inbound

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights cites this paper.

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:02:00.932716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:02:00.932716Z digest=sha256:a40ac3d4b2d1018864a7735e76ca6dbbc5c73e229b4d8751ffb0cf2c0625cc12

Observation aab1ca16-6c06-452b-b232-3451f46e26be · inbound

Data Mixing for Large Language Models Pretraining: A Survey and Outlook cites this paper.

Data Mixing for Large Language Models Pretraining: A Survey and Outlook R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.732902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T00:56:04.958757Z digest=sha256:5099b6ea1342a0f88a7c659d01d2045e63f06e873f8f57295ac4e106b1fe1ee1

Observation b81392e0-504d-4e37-90eb-d8aa1db6d013 · inbound

WARP: Weight-Space Analysis for Recovering Training Data Portfolios cites this paper.

WARP: Weight-Space Analysis for Recovering Training Data Portfolios R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:58:46.762888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-03T17:54:31.856386Z digest=sha256:386392aa784e60e7581c53ef962d4d9743cf688545a90ffe8d111acb3de5c8d3