{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:Z5QZSGALKNFAJOBCV6VRCO3ORO","short_pith_number":"pith:Z5QZSGAL","schema_version":"1.0","canonical_sha256":"cf6199180b534a04b822afab113b6e8bba3c0e6b77bcfaa9bca3d2761a66e4e2","source":{"kind":"arxiv","id":"2506.14111","version":2},"attestation_state":"computed","paper":{"title":"Essential-Web v1.0: 24T tokens of organized web data","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.LG"],"primary_cat":"cs.CL","authors_text":"Adarsh Chaluvaraju, Alok Tripathy, Anil Thomas, Ashish Tanwer, Ashish Vaswani, Darsh J Shah, Essential AI: Andrew Hojel, Ishaan Shah, Karl Stratos, Khoi Nguyen, Kurt Smith, Michael Callahan, Michael Pust, Mohit Parmar, Peter Rushton, Philip Monk, Platon Mazarakis, Ritvik Kapila, Saad Jamal, Saurabh Srivastava, Somanshu Singla, Tim Romanski, Yash Vanjani","submitted_at":"2025-06-17T02:03:36Z","abstract_excerpt":"Data plays the most prominent role in how language models acquire skills and knowledge. The lack of massive, well-organized pre-training datasets results in costly and inaccessible data pipelines. We present Essential-Web v1.0, a 24-trillion-token dataset in which every document is annotated with a twelve-category taxonomy covering topic, format, content complexity, and quality. Taxonomy labels are produced by EAI-Distill-0.5b, a fine-tuned 0.5b-parameter model that achieves an annotator agreement within 3% of Qwen2.5-32B-Instruct. With nothing more than SQL-style filters, we obtain competitiv"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2506.14111","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2025-06-17T02:03:36Z","cross_cats_sorted":["cs.AI","cs.LG"],"title_canon_sha256":"51c9f64fa186c5c25846d1ce65dc48f0f5fd8f473029d7f98083540dd546747b","abstract_canon_sha256":"3f0d330d57d09a8943732f89b57091b7b7686838193ed46b40e4b30c37ebaa44"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:24:24.232107Z","signature_b64":"z54sqzhvY7Fzn+o8uJuVqjPe4ZC5xUkeGPLej/A5C1EAGDcopn4XpfU9jmcjprgq1iMmsTep9pyw7t5J9ZyEAg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"cf6199180b534a04b822afab113b6e8bba3c0e6b77bcfaa9bca3d2761a66e4e2","last_reissued_at":"2026-07-05T11:24:24.231434Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:24:24.231434Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Essential-Web v1.0: 24T tokens of organized web data","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.LG"],"primary_cat":"cs.CL","authors_text":"Adarsh Chaluvaraju, Alok Tripathy, Anil Thomas, Ashish Tanwer, Ashish Vaswani, Darsh J Shah, Essential AI: Andrew Hojel, Ishaan Shah, Karl Stratos, Khoi Nguyen, Kurt Smith, Michael Callahan, Michael Pust, Mohit Parmar, Peter Rushton, Philip Monk, Platon Mazarakis, Ritvik Kapila, Saad Jamal, Saurabh Srivastava, Somanshu Singla, Tim Romanski, Yash Vanjani","submitted_at":"2025-06-17T02:03:36Z","abstract_excerpt":"Data plays the most prominent role in how language models acquire skills and knowledge. The lack of massive, well-organized pre-training datasets results in costly and inaccessible data pipelines. We present Essential-Web v1.0, a 24-trillion-token dataset in which every document is annotated with a twelve-category taxonomy covering topic, format, content complexity, and quality. Taxonomy labels are produced by EAI-Distill-0.5b, a fine-tuned 0.5b-parameter model that achieves an annotator agreement within 3% of Qwen2.5-32B-Instruct. With nothing more than SQL-style filters, we obtain competitiv"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2506.14111","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2506.14111/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2506.14111","created_at":"2026-07-05T11:24:24.231517+00:00"},{"alias_kind":"arxiv_version","alias_value":"2506.14111v2","created_at":"2026-07-05T11:24:24.231517+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2506.14111","created_at":"2026-07-05T11:24:24.231517+00:00"},{"alias_kind":"pith_short_12","alias_value":"Z5QZSGALKNFA","created_at":"2026-07-05T11:24:24.231517+00:00"},{"alias_kind":"pith_short_16","alias_value":"Z5QZSGALKNFAJOBC","created_at":"2026-07-05T11:24:24.231517+00:00"},{"alias_kind":"pith_short_8","alias_value":"Z5QZSGAL","created_at":"2026-07-05T11:24:24.231517+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":3,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.07778","citing_title":"Unlocking Latent Value: Taxonomy-Guided Recovery of High-Performing Data from Low-Tier Web Corpora","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12715","citing_title":"Scaling Laws for Mixture Pretraining Under Data Constraints","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12715","citing_title":"Scaling Laws for Mixture Pretraining Under Data Constraints","ref_index":16,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/Z5QZSGALKNFAJOBCV6VRCO3ORO","json":"https://pith.science/pith/Z5QZSGALKNFAJOBCV6VRCO3ORO.json","graph_json":"https://pith.science/api/pith-number/Z5QZSGALKNFAJOBCV6VRCO3ORO/graph.json","events_json":"https://pith.science/api/pith-number/Z5QZSGALKNFAJOBCV6VRCO3ORO/events.json","paper":"https://pith.science/paper/Z5QZSGAL"},"agent_actions":{"view_html":"https://pith.science/pith/Z5QZSGALKNFAJOBCV6VRCO3ORO","download_json":"https://pith.science/pith/Z5QZSGALKNFAJOBCV6VRCO3ORO.json","view_paper":"https://pith.science/paper/Z5QZSGAL","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2506.14111&json=true","fetch_graph":"https://pith.science/api/pith-number/Z5QZSGALKNFAJOBCV6VRCO3ORO/graph.json","fetch_events":"https://pith.science/api/pith-number/Z5QZSGALKNFAJOBCV6VRCO3ORO/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/Z5QZSGALKNFAJOBCV6VRCO3ORO/action/timestamp_anchor","attest_storage":"https://pith.science/pith/Z5QZSGALKNFAJOBCV6VRCO3ORO/action/storage_attestation","attest_author":"https://pith.science/pith/Z5QZSGALKNFAJOBCV6VRCO3ORO/action/author_attestation","sign_citation":"https://pith.science/pith/Z5QZSGALKNFAJOBCV6VRCO3ORO/action/citation_signature","submit_replication":"https://pith.science/pith/Z5QZSGALKNFAJOBCV6VRCO3ORO/action/replication_record"}},"created_at":"2026-07-05T11:24:24.231517+00:00","updated_at":"2026-07-05T11:24:24.231517+00:00"}