{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:TFGRUFSJMKIAYD5QZ4N4SBWZ6D","short_pith_number":"pith:TFGRUFSJ","schema_version":"1.0","canonical_sha256":"994d1a164962900c0fb0cf1bc906d9f0ec34fe3690ca2946dae217505bdd947c","source":{"kind":"arxiv","id":"2402.01306","version":4},"attestation_state":"computed","paper":{"title":"KTO: Model Alignment as Prospect Theoretic Optimization","license":"http://creativecommons.org/licenses/by-sa/4.0/","headline":"KTO aligns LLMs by maximizing prospect-theoretic utility from binary desirability signals rather than paired preferences.","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Dan Jurafsky, Douwe Kiela, Kawin Ethayarajh, Niklas Muennighoff, Winnie Xu","submitted_at":"2024-02-02T10:53:36Z","abstract_excerpt":"Kahneman & Tversky's $\\textit{prospect theory}$ tells us that humans perceive random variables in a biased but well-defined manner (1992); for example, humans are famously loss-averse. We show that objectives for aligning LLMs with human feedback implicitly incorporate many of these biases -- the success of these objectives (e.g., DPO) over cross-entropy minimization can partly be ascribed to them belonging to a family of loss functions that we call $\\textit{human-aware losses}$ (HALOs). However, the utility functions these methods attribute to humans still differ from those in the prospect th"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"2402.01306","kind":"arxiv","version":4},"metadata":{"license":"http://creativecommons.org/licenses/by-sa/4.0/","primary_cat":"cs.LG","submitted_at":"2024-02-02T10:53:36Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"8a08c1d7a8bd0d9b6fcb32559ffffe8c804c7d7c9713ba1fe5f0aa40351edd9e","abstract_canon_sha256":"94efdd36e188101bee0587c220581f370ce6a2771feaf9014f32fa1ee0a992f8"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:37:35.138979Z","signature_b64":"11pqBVvH7qwpuu1l68GQSzJfZ9aqXg3mJQSp3Ys9I6vHKPgn2+w4iHe7i846f72vnbOnG6nYM7IoLrf6yEVcDA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"994d1a164962900c0fb0cf1bc906d9f0ec34fe3690ca2946dae217505bdd947c","last_reissued_at":"2026-07-05T09:37:35.138392Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:37:35.138392Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"KTO: Model Alignment as Prospect Theoretic Optimization","license":"http://creativecommons.org/licenses/by-sa/4.0/","headline":"KTO aligns LLMs by maximizing prospect-theoretic utility from binary desirability signals rather than paired preferences.","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Dan Jurafsky, Douwe Kiela, Kawin Ethayarajh, Niklas Muennighoff, Winnie Xu","submitted_at":"2024-02-02T10:53:36Z","abstract_excerpt":"Kahneman & Tversky's $\\textit{prospect theory}$ tells us that humans perceive random variables in a biased but well-defined manner (1992); for example, humans are famously loss-averse. We show that objectives for aligning LLMs with human feedback implicitly incorporate many of these biases -- the success of these objectives (e.g., DPO) over cross-entropy minimization can partly be ascribed to them belonging to a family of loss functions that we call $\\textit{human-aware losses}$ (HALOs). However, the utility functions these methods attribute to humans still differ from those in the prospect th"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Using a Kahneman-Tversky model of human utility, we propose a HALO that directly maximizes the utility of generations instead of maximizing the log-likelihood of preferences, as current methods do. We call this approach KTO, and it matches or exceeds the performance of preference-based methods at scales from 1B to 30B, despite only learning from a binary signal of whether an output is desirable.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the specific utility function taken from prospect theory literature accurately captures human judgments of LLM outputs and that optimizing it with only binary desirability labels is sufficient without additional modeling assumptions or reference-point choices.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"KTO aligns LLMs by directly maximizing prospect-theoretic utility on binary signals and matches or exceeds preference-based methods like DPO from 1B to 30B parameters.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"KTO aligns LLMs by maximizing prospect-theoretic utility from binary desirability signals rather than paired preferences.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"8be333cd160e6013dbd6f299016d3a611cc12ca23d7edc7d71d1f3fc84c7a683"},"source":{"id":"2402.01306","kind":"arxiv","version":4},"verdict":{"id":"87514b3b-5042-4f1b-945b-7588d6420a54","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-12T12:13:01.699710Z","strongest_claim":"Using a Kahneman-Tversky model of human utility, we propose a HALO that directly maximizes the utility of generations instead of maximizing the log-likelihood of preferences, as current methods do. We call this approach KTO, and it matches or exceeds the performance of preference-based methods at scales from 1B to 30B, despite only learning from a binary signal of whether an output is desirable.","one_line_summary":"KTO aligns LLMs by directly maximizing prospect-theoretic utility on binary signals and matches or exceeds preference-based methods like DPO from 1B to 30B parameters.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the specific utility function taken from prospect theory literature accurately captures human judgments of LLM outputs and that optimizing it with only binary desirability labels is sufficient without additional modeling assumptions or reference-point choices.","pith_extraction_headline":"KTO aligns LLMs by maximizing prospect-theoretic utility from binary desirability signals rather than paired preferences."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2402.01306/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":31,"sample":[{"doi":"","year":null,"title":"Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback","work_id":"a1f2574b-a899-4713-be60-c87ba332656c","ref_index":1,"cited_arxiv_id":"2204.05862","is_internal_anchor":true},{"doi":"","year":null,"title":"Human irrationality: both bad and good for reward inference","work_id":"6361cb62-bdb5-4bc0-ae53-1762784ce3ed","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":null,"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","ref_index":3,"cited_arxiv_id":"2107.03374","is_internal_anchor":true},{"doi":"","year":null,"title":"Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models","work_id":"6b7f0773-4e99-4274-9d8e-279a1f25c5e1","ref_index":4,"cited_arxiv_id":"2401.01335","is_internal_anchor":true},{"doi":"","year":null,"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","ref_index":5,"cited_arxiv_id":"2110.14168","is_internal_anchor":true}],"resolved_work":31,"snapshot_sha256":"1af693d73e016bc199d6be0a67c000ebb3d08a7e45f265c763e6608aa7d896f1","internal_anchors":16},"formal_canon":{"evidence_count":2,"snapshot_sha256":"186826c2ffbc889bdc86f9b8dd7763c7adc4c359fdd8226b06438fe08e3e2ecb"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2402.01306","created_at":"2026-07-05T09:37:35.138454+00:00"},{"alias_kind":"arxiv_version","alias_value":"2402.01306v4","created_at":"2026-07-05T09:37:35.138454+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2402.01306","created_at":"2026-07-05T09:37:35.138454+00:00"},{"alias_kind":"pith_short_12","alias_value":"TFGRUFSJMKIA","created_at":"2026-07-05T09:37:35.138454+00:00"},{"alias_kind":"pith_short_16","alias_value":"TFGRUFSJMKIAYD5Q","created_at":"2026-07-05T09:37:35.138454+00:00"},{"alias_kind":"pith_short_8","alias_value":"TFGRUFSJ","created_at":"2026-07-05T09:37:35.138454+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":109,"internal_anchor_count":109,"sample":[{"citing_arxiv_id":"2607.06175","citing_title":"Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design","ref_index":15,"is_internal_anchor":true},{"citing_arxiv_id":"2607.05394","citing_title":"Weak-to-Strong Generalization via Direct On-Policy Distillation","ref_index":100,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24331","citing_title":"Transformer-Based Language Models Across Domain Verticals: Architectures, Applications and Critical Assessment","ref_index":35,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24004","citing_title":"Towards Spec Learning: Inference-Time Alignment from Preference Pairs","ref_index":29,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24937","citing_title":"The Hitchhiker's Guide to Agentic AI: From Foundations to Systems","ref_index":12,"is_internal_anchor":true},{"citing_arxiv_id":"2606.22276","citing_title":"Learning from Audio-Dependency Errors: Data Curation Strategies Based on Model Confusion Patterns in Audio Question Answering","ref_index":20,"is_internal_anchor":true},{"citing_arxiv_id":"2606.13006","citing_title":"Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech","ref_index":5,"is_internal_anchor":true},{"citing_arxiv_id":"2606.13227","citing_title":"PolyAlign: Conditional Human-Distribution Alignment","ref_index":47,"is_internal_anchor":true},{"citing_arxiv_id":"2606.12590","citing_title":"Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"2606.12360","citing_title":"Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal","ref_index":30,"is_internal_anchor":true},{"citing_arxiv_id":"2606.10528","citing_title":"Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output","ref_index":22,"is_internal_anchor":true},{"citing_arxiv_id":"2606.10064","citing_title":"Bittensor Agent Arenas as a Trajectory Primitive: Distilling a Shopping Agent from ShoppingBench Subnet Traces","ref_index":5,"is_internal_anchor":true},{"citing_arxiv_id":"2606.05468","citing_title":"FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization","ref_index":28,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03647","citing_title":"Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs","ref_index":37,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03994","citing_title":"SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"2606.02981","citing_title":"Predicting Inference-Time Scaling Gains from Labeled Validation-Set Output Statistics","ref_index":8,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03070","citing_title":"ASymPO: Asymmetric-Scale Policy Optimization for Asynchronous LLM Post-Training Without Behavior Information","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"2606.01561","citing_title":"S-SPPO: Semantic-Calibrated Self-Play Preference Optimization","ref_index":5,"is_internal_anchor":true},{"citing_arxiv_id":"2606.28401","citing_title":"Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs","ref_index":58,"is_internal_anchor":true},{"citing_arxiv_id":"2605.21854","citing_title":"CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models","ref_index":6,"is_internal_anchor":true},{"citing_arxiv_id":"2606.09850","citing_title":"Mechanistic Analysis of Alignment Algorithms in Language Models","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"2605.00994","citing_title":"Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives","ref_index":7,"is_internal_anchor":true},{"citing_arxiv_id":"2605.06582","citing_title":"PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization","ref_index":18,"is_internal_anchor":true},{"citing_arxiv_id":"2606.02609","citing_title":"Building Better Activation Oracles","ref_index":38,"is_internal_anchor":true},{"citing_arxiv_id":"2606.28707","citing_title":"BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards","ref_index":36,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":2,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/TFGRUFSJMKIAYD5QZ4N4SBWZ6D","json":"https://pith.science/pith/TFGRUFSJMKIAYD5QZ4N4SBWZ6D.json","graph_json":"https://pith.science/api/pith-number/TFGRUFSJMKIAYD5QZ4N4SBWZ6D/graph.json","events_json":"https://pith.science/api/pith-number/TFGRUFSJMKIAYD5QZ4N4SBWZ6D/events.json","paper":"https://pith.science/paper/TFGRUFSJ"},"agent_actions":{"view_html":"https://pith.science/pith/TFGRUFSJMKIAYD5QZ4N4SBWZ6D","download_json":"https://pith.science/pith/TFGRUFSJMKIAYD5QZ4N4SBWZ6D.json","view_paper":"https://pith.science/paper/TFGRUFSJ","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2402.01306&json=true","fetch_graph":"https://pith.science/api/pith-number/TFGRUFSJMKIAYD5QZ4N4SBWZ6D/graph.json","fetch_events":"https://pith.science/api/pith-number/TFGRUFSJMKIAYD5QZ4N4SBWZ6D/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/TFGRUFSJMKIAYD5QZ4N4SBWZ6D/action/timestamp_anchor","attest_storage":"https://pith.science/pith/TFGRUFSJMKIAYD5QZ4N4SBWZ6D/action/storage_attestation","attest_author":"https://pith.science/pith/TFGRUFSJMKIAYD5QZ4N4SBWZ6D/action/author_attestation","sign_citation":"https://pith.science/pith/TFGRUFSJMKIAYD5QZ4N4SBWZ6D/action/citation_signature","submit_replication":"https://pith.science/pith/TFGRUFSJMKIAYD5QZ4N4SBWZ6D/action/replication_record"}},"created_at":"2026-07-05T09:37:35.138454+00:00","updated_at":"2026-07-05T09:37:35.138454+00:00"}