{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:VSKFQMK2KU3HL6XYRDM5DVNDE6","short_pith_number":"pith:VSKFQMK2","schema_version":"1.0","canonical_sha256":"ac9458315a553675faf888d9d1d5a327b400e618fd4f0e6060372923ac4f4ee9","source":{"kind":"arxiv","id":"2509.08826","version":1},"attestation_state":"computed","paper":{"title":"RewardDance: Reward Scaling in Visual Generation","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Hanzhong Guo, Jie Liu, Jie Wu, Liang Li, Ming Li, Weilin Huang, Wei Liu, Xiaoxia Hou, Yan Zeng, Yu Gao, Zeyue Xue, Zilyu Ye","submitted_at":"2025-09-10T17:59:31Z","abstract_excerpt":"Reward Models (RMs) are critical for improving generation models via Reinforcement Learning (RL), yet the RM scaling paradigm in visual generation remains largely unexplored. It primarily due to fundamental limitations in existing approaches: CLIP-based RMs suffer from architectural and input modality constraints, while prevalent Bradley-Terry losses are fundamentally misaligned with the next-token prediction mechanism of Vision-Language Models (VLMs), hindering effective scaling. More critically, the RLHF optimization process is plagued by Reward Hacking issue, where models exploit flaws in t"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2509.08826","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2025-09-10T17:59:31Z","cross_cats_sorted":[],"title_canon_sha256":"23bf35f51d85c54d48dc9f5fb45f95c5eb5a7d1d6fbd5dc0a6099547595775cf","abstract_canon_sha256":"6e3e33f9973f5b6468492aa4e04534c27ebcfabb35a2c4ad77dce6f41fcc3a93"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T12:08:37.150600Z","signature_b64":"CqMhiODIutv3q0IiZQH4ORN1YTNSPbSKVdZnlN7iWCLXR7TeQJOxixn8Mxc1cAQbkuj8QRSqeH9jSe07CsaiDg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"ac9458315a553675faf888d9d1d5a327b400e618fd4f0e6060372923ac4f4ee9","last_reissued_at":"2026-07-05T12:08:37.150040Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T12:08:37.150040Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"RewardDance: Reward Scaling in Visual Generation","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Hanzhong Guo, Jie Liu, Jie Wu, Liang Li, Ming Li, Weilin Huang, Wei Liu, Xiaoxia Hou, Yan Zeng, Yu Gao, Zeyue Xue, Zilyu Ye","submitted_at":"2025-09-10T17:59:31Z","abstract_excerpt":"Reward Models (RMs) are critical for improving generation models via Reinforcement Learning (RL), yet the RM scaling paradigm in visual generation remains largely unexplored. It primarily due to fundamental limitations in existing approaches: CLIP-based RMs suffer from architectural and input modality constraints, while prevalent Bradley-Terry losses are fundamentally misaligned with the next-token prediction mechanism of Vision-Language Models (VLMs), hindering effective scaling. More critically, the RLHF optimization process is plagued by Reward Hacking issue, where models exploit flaws in t"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2509.08826","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2509.08826/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2509.08826","created_at":"2026-07-05T12:08:37.150115+00:00"},{"alias_kind":"arxiv_version","alias_value":"2509.08826v1","created_at":"2026-07-05T12:08:37.150115+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2509.08826","created_at":"2026-07-05T12:08:37.150115+00:00"},{"alias_kind":"pith_short_12","alias_value":"VSKFQMK2KU3H","created_at":"2026-07-05T12:08:37.150115+00:00"},{"alias_kind":"pith_short_16","alias_value":"VSKFQMK2KU3HL6XY","created_at":"2026-07-05T12:08:37.150115+00:00"},{"alias_kind":"pith_short_8","alias_value":"VSKFQMK2","created_at":"2026-07-05T12:08:37.150115+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":24,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.23626","citing_title":"DiT-Reward: Generative Representations for Text-to-Image Reward Modeling","ref_index":118,"is_internal_anchor":false},{"citing_arxiv_id":"2606.22918","citing_title":"Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation","ref_index":6,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09076","citing_title":"Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions","ref_index":58,"is_internal_anchor":false},{"citing_arxiv_id":"2606.02884","citing_title":"Are we really tilting? The mechanics of reward guidance in flow and diffusion models","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2606.00267","citing_title":"StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement","ref_index":118,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15055","citing_title":"DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11723","citing_title":"CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating","ref_index":68,"is_internal_anchor":false},{"citing_arxiv_id":"2606.27771","citing_title":"NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2606.00583","citing_title":"Improving Visual Representation Alignment Generation with GRPO","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2602.11146","citing_title":"Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling","ref_index":45,"is_internal_anchor":false},{"citing_arxiv_id":"2510.16888","citing_title":"Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2604.27505","citing_title":"Leveraging Verifier-Based Reinforcement Learning in Image Editing","ref_index":61,"is_internal_anchor":false},{"citing_arxiv_id":"2605.16951","citing_title":"Edit-GRPO: A Locality-Preserving Policy Optimization Framework for Image Editing","ref_index":37,"is_internal_anchor":false},{"citing_arxiv_id":"2512.04678","citing_title":"Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation","ref_index":77,"is_internal_anchor":false},{"citing_arxiv_id":"2512.13507","citing_title":"Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10983","citing_title":"TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2507.21802","citing_title":"MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE","ref_index":43,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05922","citing_title":"Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10983","citing_title":"TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11723","citing_title":"CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating","ref_index":56,"is_internal_anchor":false},{"citing_arxiv_id":"2509.20427","citing_title":"Seedream 4.0: Toward Next-generation Multimodal Image Generation","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2604.27505","citing_title":"Leveraging Verifier-Based Reinforcement Learning in Image Editing","ref_index":61,"is_internal_anchor":false},{"citing_arxiv_id":"2604.25427","citing_title":"A Systematic Post-Train Framework for Video Generation","ref_index":49,"is_internal_anchor":false},{"citing_arxiv_id":"2604.14148","citing_title":"Seedance 2.0: Advancing Video Generation for World Complexity","ref_index":21,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/VSKFQMK2KU3HL6XYRDM5DVNDE6","json":"https://pith.science/pith/VSKFQMK2KU3HL6XYRDM5DVNDE6.json","graph_json":"https://pith.science/api/pith-number/VSKFQMK2KU3HL6XYRDM5DVNDE6/graph.json","events_json":"https://pith.science/api/pith-number/VSKFQMK2KU3HL6XYRDM5DVNDE6/events.json","paper":"https://pith.science/paper/VSKFQMK2"},"agent_actions":{"view_html":"https://pith.science/pith/VSKFQMK2KU3HL6XYRDM5DVNDE6","download_json":"https://pith.science/pith/VSKFQMK2KU3HL6XYRDM5DVNDE6.json","view_paper":"https://pith.science/paper/VSKFQMK2","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2509.08826&json=true","fetch_graph":"https://pith.science/api/pith-number/VSKFQMK2KU3HL6XYRDM5DVNDE6/graph.json","fetch_events":"https://pith.science/api/pith-number/VSKFQMK2KU3HL6XYRDM5DVNDE6/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/VSKFQMK2KU3HL6XYRDM5DVNDE6/action/timestamp_anchor","attest_storage":"https://pith.science/pith/VSKFQMK2KU3HL6XYRDM5DVNDE6/action/storage_attestation","attest_author":"https://pith.science/pith/VSKFQMK2KU3HL6XYRDM5DVNDE6/action/author_attestation","sign_citation":"https://pith.science/pith/VSKFQMK2KU3HL6XYRDM5DVNDE6/action/citation_signature","submit_replication":"https://pith.science/pith/VSKFQMK2KU3HL6XYRDM5DVNDE6/action/replication_record"}},"created_at":"2026-07-05T12:08:37.150115+00:00","updated_at":"2026-07-05T12:08:37.150115+00:00"}