{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:FRBFVOKBD2EC2NWP3F3B7KZFVS","short_pith_number":"pith:FRBFVOKB","schema_version":"1.0","canonical_sha256":"2c425ab9411e882d36cfd9761fab25aca6b43a551b2067cf5a1dc9ce0b3a7387","source":{"kind":"arxiv","id":"2505.00358","version":1},"attestation_state":"computed","paper":{"title":"R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Albert Ge, Avi Trost, Frederic Sala, John Cooper, Kendall Park, Nicholas Roberts, Satya Sai Srinath Namburi GNVV, Tzu-Heng Huang, Ziyang Cai, Ziyi Chu","submitted_at":"2025-05-01T07:08:19Z","abstract_excerpt":"Data mixing strategies have successfully reduced the costs involved in training language models. While promising, such methods suffer from two flaws. First, they rely on predetermined data domains (e.g., data sources, task types), which may fail to capture critical semantic nuances, leaving performance on the table. Second, these methods scale with the number of domains in a computationally prohibitive way. We address these challenges via R&B, a framework that re-partitions training data based on semantic similarity (Regroup) to create finer-grained domains, and efficiently optimizes the data "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2505.00358","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2025-05-01T07:08:19Z","cross_cats_sorted":["cs.AI","cs.CL"],"title_canon_sha256":"8e2081ffda51dc2ec89ce8d4d87a29205b03bedbf0dcdb2616d7f7a8ae270f0b","abstract_canon_sha256":"86f481cae5f3b22f7ce4f643dfc562c32acecea2e380829283078e6a75b96261"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:57:12.171941Z","signature_b64":"UNPU9+FMgPsoyO0Ubmpl6Wp3A+V0YR39x36lPZ4YpB1MHMeIiP7EErcAIQuCrBzG2cWI/lE0bB+eibWz0eUjAA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"2c425ab9411e882d36cfd9761fab25aca6b43a551b2067cf5a1dc9ce0b3a7387","last_reissued_at":"2026-07-05T10:57:12.171447Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:57:12.171447Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Albert Ge, Avi Trost, Frederic Sala, John Cooper, Kendall Park, Nicholas Roberts, Satya Sai Srinath Namburi GNVV, Tzu-Heng Huang, Ziyang Cai, Ziyi Chu","submitted_at":"2025-05-01T07:08:19Z","abstract_excerpt":"Data mixing strategies have successfully reduced the costs involved in training language models. While promising, such methods suffer from two flaws. First, they rely on predetermined data domains (e.g., data sources, task types), which may fail to capture critical semantic nuances, leaving performance on the table. Second, these methods scale with the number of domains in a computationally prohibitive way. We address these challenges via R&B, a framework that re-partitions training data based on semantic similarity (Regroup) to create finer-grained domains, and efficiently optimizes the data "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2505.00358","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2505.00358/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2505.00358","created_at":"2026-07-05T10:57:12.171505+00:00"},{"alias_kind":"arxiv_version","alias_value":"2505.00358v1","created_at":"2026-07-05T10:57:12.171505+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2505.00358","created_at":"2026-07-05T10:57:12.171505+00:00"},{"alias_kind":"pith_short_12","alias_value":"FRBFVOKBD2EC","created_at":"2026-07-05T10:57:12.171505+00:00"},{"alias_kind":"pith_short_16","alias_value":"FRBFVOKBD2EC2NWP","created_at":"2026-07-05T10:57:12.171505+00:00"},{"alias_kind":"pith_short_8","alias_value":"FRBFVOKB","created_at":"2026-07-05T10:57:12.171505+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":2,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2607.01686","citing_title":"WARP: Weight-Space Analysis for Recovering Training Data Portfolios","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2604.16380","citing_title":"Data Mixing for Large Language Models Pretraining: A Survey and Outlook","ref_index":70,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/FRBFVOKBD2EC2NWP3F3B7KZFVS","json":"https://pith.science/pith/FRBFVOKBD2EC2NWP3F3B7KZFVS.json","graph_json":"https://pith.science/api/pith-number/FRBFVOKBD2EC2NWP3F3B7KZFVS/graph.json","events_json":"https://pith.science/api/pith-number/FRBFVOKBD2EC2NWP3F3B7KZFVS/events.json","paper":"https://pith.science/paper/FRBFVOKB"},"agent_actions":{"view_html":"https://pith.science/pith/FRBFVOKBD2EC2NWP3F3B7KZFVS","download_json":"https://pith.science/pith/FRBFVOKBD2EC2NWP3F3B7KZFVS.json","view_paper":"https://pith.science/paper/FRBFVOKB","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2505.00358&json=true","fetch_graph":"https://pith.science/api/pith-number/FRBFVOKBD2EC2NWP3F3B7KZFVS/graph.json","fetch_events":"https://pith.science/api/pith-number/FRBFVOKBD2EC2NWP3F3B7KZFVS/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/FRBFVOKBD2EC2NWP3F3B7KZFVS/action/timestamp_anchor","attest_storage":"https://pith.science/pith/FRBFVOKBD2EC2NWP3F3B7KZFVS/action/storage_attestation","attest_author":"https://pith.science/pith/FRBFVOKBD2EC2NWP3F3B7KZFVS/action/author_attestation","sign_citation":"https://pith.science/pith/FRBFVOKBD2EC2NWP3F3B7KZFVS/action/citation_signature","submit_replication":"https://pith.science/pith/FRBFVOKBD2EC2NWP3F3B7KZFVS/action/replication_record"}},"created_at":"2026-07-05T10:57:12.171505+00:00","updated_at":"2026-07-05T10:57:12.171505+00:00"}