{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:VISZ7I6ZTTASZ3WV5KZOXZDX2Y","short_pith_number":"pith:VISZ7I6Z","schema_version":"1.0","canonical_sha256":"aa259fa3d99cc12ceed5eab2ebe477d616236bce0ea0dc6ad2f8c78798c3aa38","source":{"kind":"arxiv","id":"2408.16673","version":2},"attestation_state":"computed","paper":{"title":"Preserving Diversity in Supervised Fine-Tuning of Large Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Congliang Chen, Jiancong Xiao, Ruoyu Sun, Tian Xu, Zeyu Qin, Zhi-Quan Luo, Ziniu Li","submitted_at":"2024-08-29T16:21:00Z","abstract_excerpt":"Large Language Models (LLMs) typically rely on Supervised Fine-Tuning (SFT) to specialize in downstream tasks, with the Cross Entropy (CE) loss being the de facto choice. However, CE maximizes the likelihood of observed data without accounting for alternative possibilities. As such, CE usually leads to reduced diversity in the model's outputs, which hinders further development that requires sampling to explore better responses. To address this limitation, this paper introduces a new game-theoretic formulation for SFT. In this framework, an auxiliary variable is introduced to regulate the learn"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2408.16673","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.LG","submitted_at":"2024-08-29T16:21:00Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"8dab21606965239ba0a545adaa130a669b651207df1f5fc934739273321b5d8c","abstract_canon_sha256":"78c72aa3903c8123f71dff7e26119e18f515e706ce537d715a9f2fd352b0058e"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:44:39.714854Z","signature_b64":"6hgdl0yS5QV4Lthp/4deKxnqZKDyOlb9Rq5MuMfpEjAW0L7Tijylg1sa/xwWOGA/wid+17k0N5a0oT6+crEkDw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"aa259fa3d99cc12ceed5eab2ebe477d616236bce0ea0dc6ad2f8c78798c3aa38","last_reissued_at":"2026-07-05T10:44:39.714344Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:44:39.714344Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Preserving Diversity in Supervised Fine-Tuning of Large Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Congliang Chen, Jiancong Xiao, Ruoyu Sun, Tian Xu, Zeyu Qin, Zhi-Quan Luo, Ziniu Li","submitted_at":"2024-08-29T16:21:00Z","abstract_excerpt":"Large Language Models (LLMs) typically rely on Supervised Fine-Tuning (SFT) to specialize in downstream tasks, with the Cross Entropy (CE) loss being the de facto choice. However, CE maximizes the likelihood of observed data without accounting for alternative possibilities. As such, CE usually leads to reduced diversity in the model's outputs, which hinders further development that requires sampling to explore better responses. To address this limitation, this paper introduces a new game-theoretic formulation for SFT. In this framework, an auxiliary variable is introduced to regulate the learn"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2408.16673","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2408.16673/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2408.16673","created_at":"2026-07-05T10:44:39.714405+00:00"},{"alias_kind":"arxiv_version","alias_value":"2408.16673v2","created_at":"2026-07-05T10:44:39.714405+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2408.16673","created_at":"2026-07-05T10:44:39.714405+00:00"},{"alias_kind":"pith_short_12","alias_value":"VISZ7I6ZTTAS","created_at":"2026-07-05T10:44:39.714405+00:00"},{"alias_kind":"pith_short_16","alias_value":"VISZ7I6ZTTASZ3WV","created_at":"2026-07-05T10:44:39.714405+00:00"},{"alias_kind":"pith_short_8","alias_value":"VISZ7I6Z","created_at":"2026-07-05T10:44:39.714405+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":12,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.29184","citing_title":"BaRA: Bayesian Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning","ref_index":59,"is_internal_anchor":false},{"citing_arxiv_id":"2606.00147","citing_title":"RAFT: Data Refinement and Adaptive Distillation for Domain Fine-Tuning with Alleviated Forgetting","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2508.17784","citing_title":"Proximal Supervised Fine-Tuning","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2601.12538","citing_title":"Agentic Reasoning for Large Language Models","ref_index":231,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11505","citing_title":"Selective Off-Policy Reference Tuning with Plan Guidance","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07076","citing_title":"Self-Consolidating Language Models: Continual Knowledge Incorporation from Context","ref_index":59,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11505","citing_title":"Selective Off-Policy Reference Tuning with Plan Guidance","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09995","citing_title":"Annotations Mitigate Post-Training Mode Collapse","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00195","citing_title":"Diversity in Large Language Models under Supervised Fine-Tuning","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07076","citing_title":"Self-Consolidating Language Models: Continual Knowledge Incorporation from Context","ref_index":59,"is_internal_anchor":false},{"citing_arxiv_id":"2604.14258","citing_title":"GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00195","citing_title":"Diversity in Large Language Models under Supervised Fine-Tuning","ref_index":7,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/VISZ7I6ZTTASZ3WV5KZOXZDX2Y","json":"https://pith.science/pith/VISZ7I6ZTTASZ3WV5KZOXZDX2Y.json","graph_json":"https://pith.science/api/pith-number/VISZ7I6ZTTASZ3WV5KZOXZDX2Y/graph.json","events_json":"https://pith.science/api/pith-number/VISZ7I6ZTTASZ3WV5KZOXZDX2Y/events.json","paper":"https://pith.science/paper/VISZ7I6Z"},"agent_actions":{"view_html":"https://pith.science/pith/VISZ7I6ZTTASZ3WV5KZOXZDX2Y","download_json":"https://pith.science/pith/VISZ7I6ZTTASZ3WV5KZOXZDX2Y.json","view_paper":"https://pith.science/paper/VISZ7I6Z","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2408.16673&json=true","fetch_graph":"https://pith.science/api/pith-number/VISZ7I6ZTTASZ3WV5KZOXZDX2Y/graph.json","fetch_events":"https://pith.science/api/pith-number/VISZ7I6ZTTASZ3WV5KZOXZDX2Y/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/VISZ7I6ZTTASZ3WV5KZOXZDX2Y/action/timestamp_anchor","attest_storage":"https://pith.science/pith/VISZ7I6ZTTASZ3WV5KZOXZDX2Y/action/storage_attestation","attest_author":"https://pith.science/pith/VISZ7I6ZTTASZ3WV5KZOXZDX2Y/action/author_attestation","sign_citation":"https://pith.science/pith/VISZ7I6ZTTASZ3WV5KZOXZDX2Y/action/citation_signature","submit_replication":"https://pith.science/pith/VISZ7I6ZTTASZ3WV5KZOXZDX2Y/action/replication_record"}},"created_at":"2026-07-05T10:44:39.714405+00:00","updated_at":"2026-07-05T10:44:39.714405+00:00"}