{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:J4ZCC52WGBLOHBKSLNBFCMHWUW","short_pith_number":"pith:J4ZCC52W","schema_version":"1.0","canonical_sha256":"4f322177563056e385525b425130f6a58c5565c2699216307e7649571c23cd15","source":{"kind":"arxiv","id":"2507.12856","version":2},"attestation_state":"computed","paper":{"title":"Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Chongli Qin, Jost Tobias Springenberg","submitted_at":"2025-07-17T07:26:54Z","abstract_excerpt":"Behavior Cloning (BC) on curated (or filtered) data is the predominant paradigm for supervised fine-tuning (SFT) of large language models; as well as for imitation learning of control policies. Here, we draw on a connection between this successful strategy and the theory and practice of finding optimal policies via Reinforcement Learning (RL). Building on existing literature, we clarify that SFT can be understood as maximizing a lower bound on the RL objective in a sparse reward setting. Giving support to its often observed good performance. From this viewpoint, we realize that a small modific"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2507.12856","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","primary_cat":"cs.LG","submitted_at":"2025-07-17T07:26:54Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"da9eb17cfdd54e81e71802e2132858cb226143419da72c242a0852a365530316","abstract_canon_sha256":"d719841017f4b0823a4e5039230ac6fe6977c92c453fdee62c735d1c251de442"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T12:06:16.368378Z","signature_b64":"lNa8RYowDL/DPINL7XR+OQCLM7WQvsrfpyLoCwibQdvc7cej1XWmoeyaZD+fNayqCuHU54mXUvV/d21P9nvIAw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"4f322177563056e385525b425130f6a58c5565c2699216307e7649571c23cd15","last_reissued_at":"2026-07-05T12:06:16.367857Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T12:06:16.367857Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Chongli Qin, Jost Tobias Springenberg","submitted_at":"2025-07-17T07:26:54Z","abstract_excerpt":"Behavior Cloning (BC) on curated (or filtered) data is the predominant paradigm for supervised fine-tuning (SFT) of large language models; as well as for imitation learning of control policies. Here, we draw on a connection between this successful strategy and the theory and practice of finding optimal policies via Reinforcement Learning (RL). Building on existing literature, we clarify that SFT can be understood as maximizing a lower bound on the RL objective in a sparse reward setting. Giving support to its often observed good performance. From this viewpoint, we realize that a small modific"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2507.12856","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2507.12856/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2507.12856","created_at":"2026-07-05T12:06:16.367921+00:00"},{"alias_kind":"arxiv_version","alias_value":"2507.12856v2","created_at":"2026-07-05T12:06:16.367921+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2507.12856","created_at":"2026-07-05T12:06:16.367921+00:00"},{"alias_kind":"pith_short_12","alias_value":"J4ZCC52WGBLO","created_at":"2026-07-05T12:06:16.367921+00:00"},{"alias_kind":"pith_short_16","alias_value":"J4ZCC52WGBLOHBKS","created_at":"2026-07-05T12:06:16.367921+00:00"},{"alias_kind":"pith_short_8","alias_value":"J4ZCC52W","created_at":"2026-07-05T12:06:16.367921+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":11,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.11206","citing_title":"Compatibility-Aware Dynamic Fine-Tuning for Large Language Models","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2606.24064","citing_title":"Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09396","citing_title":"PriFT: Prior-Support Guided Supervised Fine-Tuning","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09932","citing_title":"When RL Fails after SFT: Rejuvenating Model Plasticity for Robust SFT-to-RL Handoff","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2605.29303","citing_title":"Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2605.31455","citing_title":"DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2508.17784","citing_title":"Proximal Supervised Fine-Tuning","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2510.01379","citing_title":"Multi-LLM Orchestration for High-Quality Code Generation: Exploiting Complementary Model Strengths","ref_index":54,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10973","citing_title":"Rotation-Preserving Supervised Fine-Tuning","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00610","citing_title":"Decouple before Integration: Test-time Synthesis of SFT and RLVR Task Vectors","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2605.02469","citing_title":"Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent","ref_index":37,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/J4ZCC52WGBLOHBKSLNBFCMHWUW","json":"https://pith.science/pith/J4ZCC52WGBLOHBKSLNBFCMHWUW.json","graph_json":"https://pith.science/api/pith-number/J4ZCC52WGBLOHBKSLNBFCMHWUW/graph.json","events_json":"https://pith.science/api/pith-number/J4ZCC52WGBLOHBKSLNBFCMHWUW/events.json","paper":"https://pith.science/paper/J4ZCC52W"},"agent_actions":{"view_html":"https://pith.science/pith/J4ZCC52WGBLOHBKSLNBFCMHWUW","download_json":"https://pith.science/pith/J4ZCC52WGBLOHBKSLNBFCMHWUW.json","view_paper":"https://pith.science/paper/J4ZCC52W","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2507.12856&json=true","fetch_graph":"https://pith.science/api/pith-number/J4ZCC52WGBLOHBKSLNBFCMHWUW/graph.json","fetch_events":"https://pith.science/api/pith-number/J4ZCC52WGBLOHBKSLNBFCMHWUW/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/J4ZCC52WGBLOHBKSLNBFCMHWUW/action/timestamp_anchor","attest_storage":"https://pith.science/pith/J4ZCC52WGBLOHBKSLNBFCMHWUW/action/storage_attestation","attest_author":"https://pith.science/pith/J4ZCC52WGBLOHBKSLNBFCMHWUW/action/author_attestation","sign_citation":"https://pith.science/pith/J4ZCC52WGBLOHBKSLNBFCMHWUW/action/citation_signature","submit_replication":"https://pith.science/pith/J4ZCC52WGBLOHBKSLNBFCMHWUW/action/replication_record"}},"created_at":"2026-07-05T12:06:16.367921+00:00","updated_at":"2026-07-05T12:06:16.367921+00:00"}