Pith. sign in

Paper Citation Record · LEDGER

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining

As of 8 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2505.20380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20380 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:04:24.977197Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:17:27.176059Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact3
  • verified fuzzy4
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 891c50f9-c251-4812-a846-2013608f0c80 · outbound

This paper cites The pile: An 800gb dataset of diverse text for language modeling, 2020.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining The pile: An 800gb dataset of diverse text for language modeling, 2020

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.381295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.381295Z digest=sha256:2e3f618230a63a4faefb7c0e8ac4173105c0976a21f092332895dc219fe2636e

Observation a3914e3c-bc38-4ed8-a2e0-a02e4569b9f7 · outbound

This paper cites Redpajama: An open source recipe to reproduce llama training dataset, 2023.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Redpajama: An open source recipe to reproduce llama training dataset, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:26.325372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:04:22.438328Z digest=sha256:679ba350aa4c071fe292226d5b25e8ae988c1ec9ca0afcdda14492221dd20bbb

Observation b4303d09-6850-428d-9ed4-32fa963c15ef · outbound

This paper cites DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.550042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.550042Z digest=sha256:adce0ca0a46e279229dd07eb15dc13d9a3a97cd3e49ffa90005b2989a0b8118d

Observation 5845024c-8104-4c50-90dd-1a03c46a31a9 · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.629012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.629012Z digest=sha256:4bf52c15aec5167753c38c9064b2527e1549d9b0d9ce175da2a632b88e68c5f2

Observation 59c3bb8f-c3bd-49c5-977d-081837ee815c · outbound

This paper cites DoGE: Domain Reweighting with Generalization Estimation.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining DoGE: Domain Reweighting with Generalization Estimation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.784397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.784397Z digest=sha256:38cc091e23916b05d16c729e86e5278c70fe34f58d8c1eb0003d366820add2fa

Observation 4514bdcb-4843-496a-b304-2ba5fc63bd0c · outbound

This paper cites Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.901349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.901349Z digest=sha256:c399622ee9fa4c03b898f1d3815f334dfeb3517550a64345c1d2b754712d9cc7

Observation 9e77b75e-a9a1-4abf-8af6-eed623f1219b · outbound

This paper cites Dynamic Gradient Alignment for Online Data Mixing.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Dynamic Gradient Alignment for Online Data Mixing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.974581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.974581Z digest=sha256:43aaf12f4e7f20d4bd9e7b5369fc486587f9b52598f426e408fa8a50259d95ba

Observation ddf230ea-6079-4599-bf9e-7fea37ee7b7d · outbound

This paper cites Gradient Surgery for Multi-Task Learning.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Gradient Surgery for Multi-Task Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:23.075091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:23.075091Z digest=sha256:5dd29364f29094c1783758f5b63d9daaba4b9917f828b3e38ce7bf7a4b952025

Observation fc54a80d-348a-4890-8693-3559ff35ebce · outbound

This paper cites FAMO: Fast Adaptive Multitask Optimization.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining FAMO: Fast Adaptive Multitask Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:23.171525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:23.171525Z digest=sha256:b0a3e7c4297ffe1a79d2a58bfdda0696d9ae15887018819c29d9d1a727b23f3c

Observation 92c0370f-7511-416a-a66f-6593e8b22641 · outbound

This paper cites Task Weighting through Gradient Projection for Multitask Learning.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Task Weighting through Gradient Projection for Multitask Learning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:04:25.628062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:04:23.239267Z digest=sha256:58485e9953b9898f6dd1d963c321f1d94d8d62e50a016afc833b01e44df9e8a0

Observation 890473d1-52e4-421e-9104-c33da6a6e0ee · outbound

This paper cites Learning Models with Uniform Performance via Distributionally Robust Optimization.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Learning Models with Uniform Performance via Distributionally Robust Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:23.320223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:23.320223Z digest=sha256:973f29b0344b091ca8c9353d6e6679628384f208cfb6836ed6dfaeedfd9ac663

Observation 01494077-0de8-451a-9b61-b332c00f151f · outbound

This paper cites Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:23.384622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:23.384622Z digest=sha256:9705ab80462675152c51e7a11d243da36d7555cd7491c6faa1be9f05a7bf645b

Observation 2161bc10-f925-448e-8759-e77f382bd826 · outbound

This paper cites An Online Method for A Class of Distributionally Robust Optimization with Non-Convex Objectives.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining An Online Method for A Class of Distributionally Robust Optimization with Non-Convex Objectives

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:04:25.474251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:04:23.474114Z digest=sha256:3f34c4fd69a52d24eda2d018eda4e4ea9a82f37c71c7fbe8505af9a12813fd60

Observation 48ac4f27-bcf8-43ea-98d2-556ad10af6f0 · outbound

This paper cites Stochastic gradient methods for distributionally robust optimization with f-divergences.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Stochastic gradient methods for distributionally robust optimization with f-divergences

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:26.145954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:04:23.584016Z digest=sha256:84ffaee9c326158d62aec3599ec636ce0573adf7c03cace9e51ae10ae2b6cd22

Observation 6428e2da-6ee0-4c9d-a343-0a49a53c2441 · outbound

This paper cites Attention Is All You Need.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Attention Is All You Need

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:23.690591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:23.690591Z digest=sha256:990ef041fb1edbad8c377608ca7a0ea20219d826f4c28647d1e0669a6bd33d2b

Observation 8d2d5629-0069-4604-8384-21c8d0164403 · outbound

This paper cites Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:23.796108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:23.796108Z digest=sha256:9d19527f4acf948301b01c0ee1c6f6e12006434da340da16e7d497eccd7c9112

Observation b47ec271-8c07-4c8d-b297-fe2d808ee1c9 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:23.913426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:23.913426Z digest=sha256:991efdd156db3fc6b50749c4de6ca6e554a4fe422e7bd3f424ef425bfc8463c3

Observation f8c2bc0f-7a6c-4fa5-893d-9d91a140078c · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Crowdsourcing Multiple Choice Science Questions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.035893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.035893Z digest=sha256:57440c9ecf6c7bb989531407e133f0cba46425960c7a8c9c0a0cc5283c032a10

Observation 948a9799-9802-45c3-8935-abbffd40b00a · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.118899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.118899Z digest=sha256:3be7812edd1626b965d7e489b1ff4b16a6ad8aec63157e04a179154d1755f6ac

Observation 5d2d7f5b-e799-49fe-bfe4-cc6e5db02d67 · outbound

This paper cites LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.217182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.217182Z digest=sha256:e46030f9751d0eb40aa29f8a6f043fecccaf95db67c7e6380c05760815922d3b

Observation 4f8b13c0-bbd2-4f9c-bf77-18f58daf0c7d · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.262249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.262249Z digest=sha256:8a1df0c29341a99f8ac8c47702ded9d0f8ca9e9d86e4e009e3fc180b29f8680d

Observation 3f4148fc-f73e-403c-b182-5b53eb5ad82f · outbound

This paper cites W iki-40 B : Multilingual language model dataset.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining W iki-40 B : Multilingual language model dataset

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:25.971942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:04:24.331842Z digest=sha256:ba23f65231a177d0625494c92ffdbd7f4090bbab3f80e6caf3900d0a173018bb

Observation 32058b09-fae7-4ce8-bfc0-e861a5801d4b · outbound

This paper cites Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.403499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.403499Z digest=sha256:be375d23b02231cb724c32818be1d5db890d2073a09cbe973e5f2331074204c1

Observation 1e101994-f756-4a34-9730-29927a129382 · outbound

This paper cites Learning to Combine: Knowledge Aggregation for Multi-Source Domain Adaptation.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Learning to Combine: Knowledge Aggregation for Multi-Source Domain Adaptation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:04:25.253706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:04:24.473199Z digest=sha256:c6b47c83b2a706e6f052b3bbd81da0fdbb34d9e56f26ee0d2583f273ef2d412a

Observation 4f784259-f93c-48de-97e0-b32d974601ba · outbound

This paper cites Efficient Online Data Mixing For Language Model Pre-Training.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Efficient Online Data Mixing For Language Model Pre-Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.571162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.571162Z digest=sha256:808f7d9d36b9100d4dab0eee6eb1a3453200acf93974cd60f12709d0ecd1e61c

Observation f8a9d79b-2215-444b-a291-5b99e7ab11d2 · outbound

This paper cites Conflict-Averse Gradient Descent for Multi-task Learning.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Conflict-Averse Gradient Descent for Multi-task Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.642321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.642321Z digest=sha256:f119bb214adcb0e440590d5d14ed6ac9578881d53584bde184a066ebddec021c

Observation a31677a6-6ff4-49b5-a0c2-2baae3a7fd06 · outbound

This paper cites GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.716563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.716563Z digest=sha256:ee015c45859bfb8177ddcd592fb02e52f04a3bcea4085db9963fbf55d0c2ccb0

Observation 1fa80b2f-9c0c-4312-ad17-8504c1af612b · outbound

This paper cites Multi-Task Learning as a Bargaining Game.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Multi-Task Learning as a Bargaining Game

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.809400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.809400Z digest=sha256:54bed60e2074cac065b963f868b6c097eb9649b10f0c4a984ba7941a23c96463

Observation 3fbc81cf-d05b-4848-ae99-bc8f97745d6f · outbound

This paper cites Multiple-gradient descent algorithm ( MGDA ) for multiobjective optimization.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Multiple-gradient descent algorithm ( MGDA ) for multiobjective optimization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:04:25.848930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T14:04:24.912025Z digest=sha256:13fa63d6039b72432227ff0164662c2317d56b4a23cfa3e4220774a781bff7e1

Observation 131c85fc-1716-4469-97cf-1d194ba1e879 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:24.977197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:24.977197Z digest=sha256:f10cc1b4ea3b4771a7af83c955f3493e6d16fd2f59068565d456be88d8360c87

Pith citing papers

Observation c1292400-923a-4c76-bc8e-077a045b9afa · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.176059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.176059Z digest=sha256:9afe0da249785653e1eea84595ac049c73482505e14020ac6c7b316f1903f960