Pith. sign in

Paper Citation Record · LEDGER

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning

As of 8 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2607.22769.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.22769 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:19:33.724551Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3652de58-bbaa-4d06-8d5d-c3d3c0f48d4c · outbound

This paper cites Nemotron-climb: Clustering-based iterative data mixture optimization.

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning Nemotron-climb: Clustering-based iterative data mixture optimization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:32.099580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:32.099580Z digest=sha256:c84345451412732956e42e197d6c94aacc43fd405bac1fd3b3bd9464f25d1275

Observation ac5b3502-6737-4d99-b71a-6e35bc680979 · outbound

This paper cites Training Compute-Optimal Large Language Models.

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning Training Compute-Optimal Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:32.223883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:32.223883Z digest=sha256:d5b86e16a4a3e949560b355ae023bd105946d7d6efc1929ed624ae91c331d3a7

Observation 70f937f4-19c3-4b23-9a72-22cd41637be1 · outbound

This paper cites Scaling Laws for Neural Language Models.

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning Scaling Laws for Neural Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:32.464859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:32.464859Z digest=sha256:4613056e4d558ec1cd31d0d8889839f1f17171de61cc13448c108a748005a421

Observation eef586f3-1521-4934-b262-7cb4f560ebc5 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning Rho-1: Not All Tokens Are What You Need

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:32.644746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:32.644746Z digest=sha256:231db3ad10bd4cabfd21b54d1c872ad6fd4c9e917616a620c5a0425cd019eb94

Observation f7098e97-af5e-4fc8-a46f-93f300adc2d8 · outbound

This paper cites Scalebio: Scalable bilevel optimization for llm data reweighting.Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 2025.

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning Scalebio: Scalable bilevel optimization for llm data reweighting.Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:32.754745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:32.754745Z digest=sha256:5a138d5b8f4449599e43fe6d70824afdd231cf072c876ab74316d039e445d55a

Observation c45f3513-2912-4ade-820f-f20ce1a27d97 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:32.909831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:32.909831Z digest=sha256:d30f96e8681b0152887841f0c1226fb1a064e7a456f1792d8e76a3d1442f006a

Observation 875e4b27-41d0-41de-a15c-9e8403d60a23 · outbound

This paper cites Qurating: Select- ing high-quality data for training language models.International Conference on Machine Learning, 2024.

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning Qurating: Select- ing high-quality data for training language models.International Conference on Machine Learning, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:33.024740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:33.024740Z digest=sha256:87b217c7e26515a362704221f5f15fbe9e4adb4b8b046ca16a5040341316985d

Observation c2c8288b-7575-47fa-8bed-2d1642d7603f · outbound

This paper cites LESS: Selecting Influential Data for Targeted Instruction Tuning.

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning LESS: Selecting Influential Data for Targeted Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:33.174874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:33.174874Z digest=sha256:e1cbdd4b7e8a0fa37dc6a9e1744a7d8cf5fd977bdc650da0d2ccc2603db3bfe9

Observation 24290403-1671-403d-af9b-c749768a8733 · outbound

This paper cites Data selection for language models via importance resampled mcmc.Advances in Neural Information Processing Systems, 36, 2023.

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning Data selection for language models via importance resampled mcmc.Advances in Neural Information Processing Systems, 36, 2023

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:33.305650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:33.305650Z digest=sha256:173666aa246b97674a69b7905ccd8ba99de680e3a83bf37a46a2ef114f8a4d7a

Observation 5915e398-61dc-40ce-88ee-7d05e39ceffd · outbound

This paper cites Doremi: Optimizing data mixtures speeds up language model pretraining.Advances in Neural Information Processing Systems, 36, 2023.

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning Doremi: Optimizing data mixtures speeds up language model pretraining.Advances in Neural Information Processing Systems, 36, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:33.392135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:33.392135Z digest=sha256:1148a5240d6fdb53eb85f071825d9b55ca3ff7e8f2a03f24d1c6fd2e79c7cbc2

Observation cffd2bca-5698-486b-95d8-a2832d153890 · outbound

This paper cites Dataflex: A unified framework for data-centric dynamic training of large language models.arXiv preprint arXiv:2603.26164, 2026.

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning Dataflex: A unified framework for data-centric dynamic training of large language models.arXiv preprint arXiv:2603.26164, 2026

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:33.624844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:33.624844Z digest=sha256:d4a2df2453e38e46c48f4a54a4024f45efa04691cf56b5bd142f54896aa9d017

Observation 3478ff6d-deba-4dae-bcd6-0b7517a56165 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T06:19:33.724551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:19:33.724551Z digest=sha256:dcfddef17fae89f338115e8f8c5350d7ca45a87251d6a3add9217331bbd43e5d

Pith citing papers

No inbound Pith citation observations are available.