{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:KWI5AGNIKCNRMJETQ2GRLNB3KK","short_pith_number":"pith:KWI5AGNI","schema_version":"1.0","canonical_sha256":"5591d019a8509b162493868d15b43b52be70b9165e4b194e6793c92bf350b8f1","source":{"kind":"arxiv","id":"2502.14907","version":2},"attestation_state":"computed","paper":{"title":"GneissWeb: Preparing High Quality Data for LLMs at Scale","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Abdulhamid Adebayo, Alexei Karve, Bishwaranjan Bhattacharjee, Boris Lublinsky, Changchang Liu, Constantin Adam, David Wood, Farhan Ahmed, Hajar Emami Gohari, Herbert Woisetschl\\\"ager, Issei Yoshida, Kun-Lung Wu, Maroun Touma, Nathalie Baracaldo Angel, Nirmit Desai, Pablo Pesce, Petros Zerfos, Praneet Adusumilli, Ran Iwamoto, Revital Eres, Santosh Subhashrao Borse, Shalisha Witherspoon, Shiqiang Wang, Swanand Ravindra Kadhe, Syed Yousaf Shah, Syed Zawad, Takuyo Ohko, Wei-Han Lee, Xuan-Hong Dang, Yan Koyfman, Yi Zhou, Yuan-Chi Chang","submitted_at":"2025-02-19T00:14:29Z","abstract_excerpt":"Data quantity and quality play a vital role in determining the performance of Large Language Models (LLMs). High-quality data, in particular, can significantly boost the LLM's ability to generalize on a wide range of downstream tasks. Large pre-training datasets for leading LLMs remain inaccessible to the public, whereas many open datasets are small in size (less than 5 trillion tokens), limiting their suitability for training large models.\n  In this paper, we introduce GneissWeb, a large dataset yielding around 10 trillion tokens that caters to the data quality and quantity requirements of tr"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2502.14907","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2025-02-19T00:14:29Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"d7d1d41ad396ea1b6c1fcada0fa6b7506421e117f6b0274767367602cce16294","abstract_canon_sha256":"02416bd34178ae577bdbf935b05bc0cbb22dea6b6d4c8ecb2518b9cdcbe7f96b"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:45:14.470619Z","signature_b64":"xBmPYJR/eMr7WfDpibdyeamlGmfCGRrYqSEO6T7I1b/UmtDxcrbq9+PvUe57MtsqeAH7OCfEM8FjafqE0FzXDw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"5591d019a8509b162493868d15b43b52be70b9165e4b194e6793c92bf350b8f1","last_reissued_at":"2026-07-05T11:45:14.470137Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:45:14.470137Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"GneissWeb: Preparing High Quality Data for LLMs at Scale","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Abdulhamid Adebayo, Alexei Karve, Bishwaranjan Bhattacharjee, Boris Lublinsky, Changchang Liu, Constantin Adam, David Wood, Farhan Ahmed, Hajar Emami Gohari, Herbert Woisetschl\\\"ager, Issei Yoshida, Kun-Lung Wu, Maroun Touma, Nathalie Baracaldo Angel, Nirmit Desai, Pablo Pesce, Petros Zerfos, Praneet Adusumilli, Ran Iwamoto, Revital Eres, Santosh Subhashrao Borse, Shalisha Witherspoon, Shiqiang Wang, Swanand Ravindra Kadhe, Syed Yousaf Shah, Syed Zawad, Takuyo Ohko, Wei-Han Lee, Xuan-Hong Dang, Yan Koyfman, Yi Zhou, Yuan-Chi Chang","submitted_at":"2025-02-19T00:14:29Z","abstract_excerpt":"Data quantity and quality play a vital role in determining the performance of Large Language Models (LLMs). High-quality data, in particular, can significantly boost the LLM's ability to generalize on a wide range of downstream tasks. Large pre-training datasets for leading LLMs remain inaccessible to the public, whereas many open datasets are small in size (less than 5 trillion tokens), limiting their suitability for training large models.\n  In this paper, we introduce GneissWeb, a large dataset yielding around 10 trillion tokens that caters to the data quality and quantity requirements of tr"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2502.14907","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2502.14907/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2502.14907","created_at":"2026-07-05T11:45:14.470193+00:00"},{"alias_kind":"arxiv_version","alias_value":"2502.14907v2","created_at":"2026-07-05T11:45:14.470193+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2502.14907","created_at":"2026-07-05T11:45:14.470193+00:00"},{"alias_kind":"pith_short_12","alias_value":"KWI5AGNIKCNR","created_at":"2026-07-05T11:45:14.470193+00:00"},{"alias_kind":"pith_short_16","alias_value":"KWI5AGNIKCNRMJET","created_at":"2026-07-05T11:45:14.470193+00:00"},{"alias_kind":"pith_short_8","alias_value":"KWI5AGNI","created_at":"2026-07-05T11:45:14.470193+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":1,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2505.08971","citing_title":"Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training","ref_index":29,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/KWI5AGNIKCNRMJETQ2GRLNB3KK","json":"https://pith.science/pith/KWI5AGNIKCNRMJETQ2GRLNB3KK.json","graph_json":"https://pith.science/api/pith-number/KWI5AGNIKCNRMJETQ2GRLNB3KK/graph.json","events_json":"https://pith.science/api/pith-number/KWI5AGNIKCNRMJETQ2GRLNB3KK/events.json","paper":"https://pith.science/paper/KWI5AGNI"},"agent_actions":{"view_html":"https://pith.science/pith/KWI5AGNIKCNRMJETQ2GRLNB3KK","download_json":"https://pith.science/pith/KWI5AGNIKCNRMJETQ2GRLNB3KK.json","view_paper":"https://pith.science/paper/KWI5AGNI","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2502.14907&json=true","fetch_graph":"https://pith.science/api/pith-number/KWI5AGNIKCNRMJETQ2GRLNB3KK/graph.json","fetch_events":"https://pith.science/api/pith-number/KWI5AGNIKCNRMJETQ2GRLNB3KK/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/KWI5AGNIKCNRMJETQ2GRLNB3KK/action/timestamp_anchor","attest_storage":"https://pith.science/pith/KWI5AGNIKCNRMJETQ2GRLNB3KK/action/storage_attestation","attest_author":"https://pith.science/pith/KWI5AGNIKCNRMJETQ2GRLNB3KK/action/author_attestation","sign_citation":"https://pith.science/pith/KWI5AGNIKCNRMJETQ2GRLNB3KK/action/citation_signature","submit_replication":"https://pith.science/pith/KWI5AGNIKCNRMJETQ2GRLNB3KK/action/replication_record"}},"created_at":"2026-07-05T11:45:14.470193+00:00","updated_at":"2026-07-05T11:45:14.470193+00:00"}