{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:UPMMNFR4XSGQ4EX2VEVYTIPZZJ","short_pith_number":"pith:UPMMNFR4","schema_version":"1.0","canonical_sha256":"a3d8c6963cbc8d0e12faa92b89a1f9ca4bc752abab134662d14e37bb4ce6d50b","source":{"kind":"arxiv","id":"2407.06645","version":3},"attestation_state":"computed","paper":{"title":"Entropy Law: The Story Behind Data Compression and LLM Performance","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.LG","authors_text":"Chuhan Wu, Defu Lian, Enhong Chen, Hao Wang, Mingjia Yin, Ruiming Tang, Wei Guo, Yasheng Wang, Yong Liu, Yufei Wang","submitted_at":"2024-07-09T08:14:29Z","abstract_excerpt":"Data is the cornerstone of large language models (LLMs), but not all data is useful for model learning. Carefully selected data can better elicit the capabilities of LLMs with much less computational overhead. Most methods concentrate on evaluating the quality of individual samples in data selection, while the combinatorial effects among samples are neglected. Even if each sample is of perfect quality, their combinations may be suboptimal in teaching LLMs due to their intrinsic homogeneity or contradiction. In this paper, we aim to uncover the underlying relationships between LLM performance a"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2407.06645","kind":"arxiv","version":3},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2024-07-09T08:14:29Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"f2b1f979d3bb76baf45741ddfc4ad34850125fc7da7a5f2222471ce96f4d8d48","abstract_canon_sha256":"820e2146e740d1babc265d6b312b3a1d87ac443145e40989e7cbe08c6b4784c7"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:42:38.837286Z","signature_b64":"lWLPV7fj+ZlOuD/c/3O4PPt1sgoWSqhQFAV/MF1FufU+nfjbYopG67LT6g8etAHIUvwIq4aodL0CT5U0XnnECw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"a3d8c6963cbc8d0e12faa92b89a1f9ca4bc752abab134662d14e37bb4ce6d50b","last_reissued_at":"2026-07-05T08:42:38.836802Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:42:38.836802Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Entropy Law: The Story Behind Data Compression and LLM Performance","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.LG","authors_text":"Chuhan Wu, Defu Lian, Enhong Chen, Hao Wang, Mingjia Yin, Ruiming Tang, Wei Guo, Yasheng Wang, Yong Liu, Yufei Wang","submitted_at":"2024-07-09T08:14:29Z","abstract_excerpt":"Data is the cornerstone of large language models (LLMs), but not all data is useful for model learning. Carefully selected data can better elicit the capabilities of LLMs with much less computational overhead. Most methods concentrate on evaluating the quality of individual samples in data selection, while the combinatorial effects among samples are neglected. Even if each sample is of perfect quality, their combinations may be suboptimal in teaching LLMs due to their intrinsic homogeneity or contradiction. In this paper, we aim to uncover the underlying relationships between LLM performance a"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2407.06645","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2407.06645/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2407.06645","created_at":"2026-07-05T08:42:38.836858+00:00"},{"alias_kind":"arxiv_version","alias_value":"2407.06645v3","created_at":"2026-07-05T08:42:38.836858+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2407.06645","created_at":"2026-07-05T08:42:38.836858+00:00"},{"alias_kind":"pith_short_12","alias_value":"UPMMNFR4XSGQ","created_at":"2026-07-05T08:42:38.836858+00:00"},{"alias_kind":"pith_short_16","alias_value":"UPMMNFR4XSGQ4EX2","created_at":"2026-07-05T08:42:38.836858+00:00"},{"alias_kind":"pith_short_8","alias_value":"UPMMNFR4","created_at":"2026-07-05T08:42:38.836858+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":6,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.28589","citing_title":"Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2606.05868","citing_title":"YouZhi: Towards High-Concurrency Financial LLMs via Adaptive GQA-to-MLA Transition","ref_index":61,"is_internal_anchor":false},{"citing_arxiv_id":"2606.28589","citing_title":"Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2508.04149","citing_title":"Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap","ref_index":50,"is_internal_anchor":false},{"citing_arxiv_id":"2509.22343","citing_title":"Transformers Can Learn Connectivity in Some Graphs but Not Others","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2510.18900","citing_title":"Foundation Models for Discovery and Exploration in Chemical Space","ref_index":285,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/UPMMNFR4XSGQ4EX2VEVYTIPZZJ","json":"https://pith.science/pith/UPMMNFR4XSGQ4EX2VEVYTIPZZJ.json","graph_json":"https://pith.science/api/pith-number/UPMMNFR4XSGQ4EX2VEVYTIPZZJ/graph.json","events_json":"https://pith.science/api/pith-number/UPMMNFR4XSGQ4EX2VEVYTIPZZJ/events.json","paper":"https://pith.science/paper/UPMMNFR4"},"agent_actions":{"view_html":"https://pith.science/pith/UPMMNFR4XSGQ4EX2VEVYTIPZZJ","download_json":"https://pith.science/pith/UPMMNFR4XSGQ4EX2VEVYTIPZZJ.json","view_paper":"https://pith.science/paper/UPMMNFR4","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2407.06645&json=true","fetch_graph":"https://pith.science/api/pith-number/UPMMNFR4XSGQ4EX2VEVYTIPZZJ/graph.json","fetch_events":"https://pith.science/api/pith-number/UPMMNFR4XSGQ4EX2VEVYTIPZZJ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/UPMMNFR4XSGQ4EX2VEVYTIPZZJ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/UPMMNFR4XSGQ4EX2VEVYTIPZZJ/action/storage_attestation","attest_author":"https://pith.science/pith/UPMMNFR4XSGQ4EX2VEVYTIPZZJ/action/author_attestation","sign_citation":"https://pith.science/pith/UPMMNFR4XSGQ4EX2VEVYTIPZZJ/action/citation_signature","submit_replication":"https://pith.science/pith/UPMMNFR4XSGQ4EX2VEVYTIPZZJ/action/replication_record"}},"created_at":"2026-07-05T08:42:38.836858+00:00","updated_at":"2026-07-05T08:42:38.836858+00:00"}