Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:04:24.977197Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2505.20380.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:04:24.977197Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T04:17:27.176059Z
A source-named dated measurement, never combined with another source.
Source: cited_works
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 891c50f9-c251-4812-a846-2013608f0c80 · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining The pile: An 800gb dataset of diverse text for language modeling, 2020
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3914e3c-bc38-4ed8-a2e0-a02e4569b9f7 · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Redpajama: An open source recipe to reproduce llama training dataset, 2023
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4303d09-6850-428d-9ed4-32fa963c15ef · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5845024c-8104-4c50-90dd-1a03c46a31a9 · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining RegMix: Data Mixture as Regression for Language Model Pre-training
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59c3bb8f-c3bd-49c5-977d-081837ee815c · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining DoGE: Domain Reweighting with Generalization Estimation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4514bdcb-4843-496a-b304-2ba5fc63bd0c · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e77b75e-a9a1-4abf-8af6-eed623f1219b · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Dynamic Gradient Alignment for Online Data Mixing
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddf230ea-6079-4599-bf9e-7fea37ee7b7d · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Gradient Surgery for Multi-Task Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc54a80d-348a-4890-8693-3559ff35ebce · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining FAMO: Fast Adaptive Multitask Optimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92c0370f-7511-416a-a66f-6593e8b22641 · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Task Weighting through Gradient Projection for Multitask Learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 890473d1-52e4-421e-9104-c33da6a6e0ee · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Learning Models with Uniform Performance via Distributionally Robust Optimization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01494077-0de8-451a-9b61-b332c00f151f · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2161bc10-f925-448e-8759-e77f382bd826 · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining An Online Method for A Class of Distributionally Robust Optimization with Non-Convex Objectives
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48ac4f27-bcf8-43ea-98d2-556ad10af6f0 · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Stochastic gradient methods for distributionally robust optimization with f-divergences
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6428e2da-6ee0-4c9d-a343-0a49a53c2441 · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Attention Is All You Need
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d2d5629-0069-4604-8384-21c8d0164403 · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b47ec271-8c07-4c8d-b297-fe2d808ee1c9 · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8c2bc0f-7a6c-4fa5-893d-9d91a140078c · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Crowdsourcing Multiple Choice Science Questions
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 948a9799-9802-45c3-8935-abbffd40b00a · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining PIQA: Reasoning about Physical Commonsense in Natural Language
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d2d7f5b-e799-49fe-bfe4-cc6e5db02d67 · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f8b13c0-bbd2-4f9c-bf77-18f58daf0c7d · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f4148fc-f73e-403c-b182-5b53eb5ad82f · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining W iki-40 B : Multilingual language model dataset
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32058b09-fae7-4ce8-bfc0-e861a5801d4b · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e101994-f756-4a34-9730-29927a129382 · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Learning to Combine: Knowledge Aggregation for Multi-Source Domain Adaptation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4f784259-f93c-48de-97e0-b32d974601ba · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Efficient Online Data Mixing For Language Model Pre-Training
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8a9d79b-2215-444b-a291-5b99e7ab11d2 · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Conflict-Averse Gradient Descent for Multi-task Learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a31677a6-6ff4-49b5-a0c2-2baae3a7fd06 · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fa80b2f-9c0c-4312-ad17-8504c1af612b · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Multi-Task Learning as a Bargaining Game
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fbc81cf-d05b-4848-ae99-bc8f97745d6f · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining Multiple-gradient descent algorithm ( MGDA ) for multiobjective optimization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 131c85fc-1716-4469-97cf-1d194ba1e879 · outbound
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1292400-923a-4c76-bc8e-077a045b9afa · inbound
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.