{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:TDJXDWZLTSGGGYW57YCLKO2JLG","short_pith_number":"pith:TDJXDWZL","schema_version":"1.0","canonical_sha256":"98d371db2b9c8c6362ddfe04b53b4959a10258847448b097e037aea225e26c2c","source":{"kind":"arxiv","id":"2405.13954","version":1},"attestation_state":"computed","paper":{"title":"What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Adithya Pratapa, Eduard Hovy, Emma Strubell, Eric Xing, Hwijeen Ahn, Jeff Schneider, Juhan Bae, Kewen Zhao, Minsoo Kang, Roger Grosse, Sang Keun Choe, Teruko Mitamura, Willie Neiswanger, Youngseog Chung","submitted_at":"2024-05-22T19:39:05Z","abstract_excerpt":"Large language models (LLMs) are trained on a vast amount of human-written data, but data providers often remain uncredited. In response to this issue, data valuation (or data attribution), which quantifies the contribution or value of each data to the model output, has been discussed as a potential solution. Nevertheless, applying existing data valuation methods to recent LLMs and their vast training datasets has been largely limited by prohibitive compute and memory costs. In this work, we focus on influence functions, a popular gradient-based data valuation method, and significantly improve"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2405.13954","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2024-05-22T19:39:05Z","cross_cats_sorted":["cs.AI","cs.CL"],"title_canon_sha256":"b86c58f35ced1f34ad8c6b0b137f0bf836f10775a16dee8d0263c01fe170d977","abstract_canon_sha256":"83212063679467943bf47fd74bcaca59f422706cf75e4f6d27846d3bd1a5d5d1"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:22:08.736516Z","signature_b64":"uy9/kZezyCSdyWg9/UTWIVwyZvQpKK10M08INybyZ7L8f5gu6BAXjDp+x/7MVdt9Pd4LO/wTiMpYK0pyDVU3Aw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"98d371db2b9c8c6362ddfe04b53b4959a10258847448b097e037aea225e26c2c","last_reissued_at":"2026-07-05T08:22:08.735981Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:22:08.735981Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.LG","authors_text":"Adithya Pratapa, Eduard Hovy, Emma Strubell, Eric Xing, Hwijeen Ahn, Jeff Schneider, Juhan Bae, Kewen Zhao, Minsoo Kang, Roger Grosse, Sang Keun Choe, Teruko Mitamura, Willie Neiswanger, Youngseog Chung","submitted_at":"2024-05-22T19:39:05Z","abstract_excerpt":"Large language models (LLMs) are trained on a vast amount of human-written data, but data providers often remain uncredited. In response to this issue, data valuation (or data attribution), which quantifies the contribution or value of each data to the model output, has been discussed as a potential solution. Nevertheless, applying existing data valuation methods to recent LLMs and their vast training datasets has been largely limited by prohibitive compute and memory costs. In this work, we focus on influence functions, a popular gradient-based data valuation method, and significantly improve"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2405.13954","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2405.13954/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2405.13954","created_at":"2026-07-05T08:22:08.736045+00:00"},{"alias_kind":"arxiv_version","alias_value":"2405.13954v1","created_at":"2026-07-05T08:22:08.736045+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2405.13954","created_at":"2026-07-05T08:22:08.736045+00:00"},{"alias_kind":"pith_short_12","alias_value":"TDJXDWZLTSGG","created_at":"2026-07-05T08:22:08.736045+00:00"},{"alias_kind":"pith_short_16","alias_value":"TDJXDWZLTSGGGYW5","created_at":"2026-07-05T08:22:08.736045+00:00"},{"alias_kind":"pith_short_8","alias_value":"TDJXDWZL","created_at":"2026-07-05T08:22:08.736045+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":11,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.18307","citing_title":"DRIFT: Refining Instruction Data via On-Policy Data Attribution","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2606.11616","citing_title":"DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence Vectors","ref_index":5,"is_internal_anchor":false},{"citing_arxiv_id":"2607.00641","citing_title":"What's a Credit Worth? A Market Framework for Attribution-Aware Compensation in Generative Music","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2606.02142","citing_title":"TimeBlocks: Foundational and Continual Time-Series Blockbase -- Extended Version","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2606.23591","citing_title":"Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2412.08637","citing_title":"DMin: Scalable Training Data Influence Estimation for Diffusion Models","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2512.12572","citing_title":"On the Accuracy of Newton Step and Influence Function Data Attributions","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2602.03783","citing_title":"Efficient Estimation of Kernel Surrogate Models for Task Attribution","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2602.10995","citing_title":"A Human-Centric Framework for Data Attribution in Large Language Models","ref_index":50,"is_internal_anchor":false},{"citing_arxiv_id":"2603.19297","citing_title":"CLaRE-ty Amid Chaos: Quantifying Representational Entanglement to Predict Ripple Effects in LLM Editing","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2604.16197","citing_title":"Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation","ref_index":10,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/TDJXDWZLTSGGGYW57YCLKO2JLG","json":"https://pith.science/pith/TDJXDWZLTSGGGYW57YCLKO2JLG.json","graph_json":"https://pith.science/api/pith-number/TDJXDWZLTSGGGYW57YCLKO2JLG/graph.json","events_json":"https://pith.science/api/pith-number/TDJXDWZLTSGGGYW57YCLKO2JLG/events.json","paper":"https://pith.science/paper/TDJXDWZL"},"agent_actions":{"view_html":"https://pith.science/pith/TDJXDWZLTSGGGYW57YCLKO2JLG","download_json":"https://pith.science/pith/TDJXDWZLTSGGGYW57YCLKO2JLG.json","view_paper":"https://pith.science/paper/TDJXDWZL","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2405.13954&json=true","fetch_graph":"https://pith.science/api/pith-number/TDJXDWZLTSGGGYW57YCLKO2JLG/graph.json","fetch_events":"https://pith.science/api/pith-number/TDJXDWZLTSGGGYW57YCLKO2JLG/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/TDJXDWZLTSGGGYW57YCLKO2JLG/action/timestamp_anchor","attest_storage":"https://pith.science/pith/TDJXDWZLTSGGGYW57YCLKO2JLG/action/storage_attestation","attest_author":"https://pith.science/pith/TDJXDWZLTSGGGYW57YCLKO2JLG/action/author_attestation","sign_citation":"https://pith.science/pith/TDJXDWZLTSGGGYW57YCLKO2JLG/action/citation_signature","submit_replication":"https://pith.science/pith/TDJXDWZLTSGGGYW57YCLKO2JLG/action/replication_record"}},"created_at":"2026-07-05T08:22:08.736045+00:00","updated_at":"2026-07-05T08:22:08.736045+00:00"}