{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:XJZ2HTY44OCAQCGGGKGTOC2IT7","short_pith_number":"pith:XJZ2HTY4","schema_version":"1.0","canonical_sha256":"ba73a3cf1ce3840808c6328d370b489fe74c193f488916a95cc4d4109d8e404b","source":{"kind":"arxiv","id":"2505.18536","version":1},"attestation_state":"computed","paper":{"title":"Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CV"],"primary_cat":"cs.CL","authors_text":"Bo Xia, Haoyuan Sun, Jiaqi Wu, Kai Qin, Tiantian Zhang, Xueqian Wang, Xufei Lv, Yifei Zhao, Yifu Luo, Yongzhe Chang","submitted_at":"2025-05-24T06:01:48Z","abstract_excerpt":"Standing in 2025, at a critical juncture in the pursuit of Artificial General Intelligence (AGI), reinforcement fine-tuning (RFT) has demonstrated significant potential in enhancing the reasoning capability of large language models (LLMs) and has led to the development of cutting-edge AI models such as OpenAI-o1 and DeepSeek-R1. Moreover, the efficient application of RFT to enhance the reasoning capability of multimodal large language models (MLLMs) has attracted widespread attention from the community. In this position paper, we argue that reinforcement fine-tuning powers the reasoning capabi"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2505.18536","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2025-05-24T06:01:48Z","cross_cats_sorted":["cs.AI","cs.CV"],"title_canon_sha256":"28588b888a268492895d44446dc00817533d638afdea7bc2b11a8ef74bf45d7d","abstract_canon_sha256":"ed2b99eacd38f05f783283688172e0fd637f87cdf6d1fc625651589d0b6ccfa3"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:08:47.965675Z","signature_b64":"lLQ7YJA9trNnOfCuAKRcCXyBRndZFPr9fJ5hP+1ss9ROkxicLLS+MIzc3y6uWUnSknAKw+7p3FjpOhd1DCQgAQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"ba73a3cf1ce3840808c6328d370b489fe74c193f488916a95cc4d4109d8e404b","last_reissued_at":"2026-07-05T11:08:47.965107Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:08:47.965107Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CV"],"primary_cat":"cs.CL","authors_text":"Bo Xia, Haoyuan Sun, Jiaqi Wu, Kai Qin, Tiantian Zhang, Xueqian Wang, Xufei Lv, Yifei Zhao, Yifu Luo, Yongzhe Chang","submitted_at":"2025-05-24T06:01:48Z","abstract_excerpt":"Standing in 2025, at a critical juncture in the pursuit of Artificial General Intelligence (AGI), reinforcement fine-tuning (RFT) has demonstrated significant potential in enhancing the reasoning capability of large language models (LLMs) and has led to the development of cutting-edge AI models such as OpenAI-o1 and DeepSeek-R1. Moreover, the efficient application of RFT to enhance the reasoning capability of multimodal large language models (MLLMs) has attracted widespread attention from the community. In this position paper, we argue that reinforcement fine-tuning powers the reasoning capabi"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2505.18536","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2505.18536/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2505.18536","created_at":"2026-07-05T11:08:47.965181+00:00"},{"alias_kind":"arxiv_version","alias_value":"2505.18536v1","created_at":"2026-07-05T11:08:47.965181+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2505.18536","created_at":"2026-07-05T11:08:47.965181+00:00"},{"alias_kind":"pith_short_12","alias_value":"XJZ2HTY44OCA","created_at":"2026-07-05T11:08:47.965181+00:00"},{"alias_kind":"pith_short_16","alias_value":"XJZ2HTY44OCAQCGG","created_at":"2026-07-05T11:08:47.965181+00:00"},{"alias_kind":"pith_short_8","alias_value":"XJZ2HTY4","created_at":"2026-07-05T11:08:47.965181+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":8,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.25634","citing_title":"SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity","ref_index":56,"is_internal_anchor":false},{"citing_arxiv_id":"2606.21498","citing_title":"Balancing Performance and Diversity in GRPO Autoregressive Text-to-Image Post-Training","ref_index":28,"is_internal_anchor":false},{"citing_arxiv_id":"2511.20785","citing_title":"LongVT: Incentivizing \"Thinking with Long Videos\" via Native Tool Calling","ref_index":40,"is_internal_anchor":false},{"citing_arxiv_id":"2510.21583","citing_title":"Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2512.03043","citing_title":"OneThinker: All-in-one Reasoning Model for Image and Video","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2602.00181","citing_title":"CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning","ref_index":41,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10937","citing_title":"Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping","ref_index":46,"is_internal_anchor":false},{"citing_arxiv_id":"2604.08539","citing_title":"OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks","ref_index":31,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/XJZ2HTY44OCAQCGGGKGTOC2IT7","json":"https://pith.science/pith/XJZ2HTY44OCAQCGGGKGTOC2IT7.json","graph_json":"https://pith.science/api/pith-number/XJZ2HTY44OCAQCGGGKGTOC2IT7/graph.json","events_json":"https://pith.science/api/pith-number/XJZ2HTY44OCAQCGGGKGTOC2IT7/events.json","paper":"https://pith.science/paper/XJZ2HTY4"},"agent_actions":{"view_html":"https://pith.science/pith/XJZ2HTY44OCAQCGGGKGTOC2IT7","download_json":"https://pith.science/pith/XJZ2HTY44OCAQCGGGKGTOC2IT7.json","view_paper":"https://pith.science/paper/XJZ2HTY4","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2505.18536&json=true","fetch_graph":"https://pith.science/api/pith-number/XJZ2HTY44OCAQCGGGKGTOC2IT7/graph.json","fetch_events":"https://pith.science/api/pith-number/XJZ2HTY44OCAQCGGGKGTOC2IT7/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/XJZ2HTY44OCAQCGGGKGTOC2IT7/action/timestamp_anchor","attest_storage":"https://pith.science/pith/XJZ2HTY44OCAQCGGGKGTOC2IT7/action/storage_attestation","attest_author":"https://pith.science/pith/XJZ2HTY44OCAQCGGGKGTOC2IT7/action/author_attestation","sign_citation":"https://pith.science/pith/XJZ2HTY44OCAQCGGGKGTOC2IT7/action/citation_signature","submit_replication":"https://pith.science/pith/XJZ2HTY44OCAQCGGGKGTOC2IT7/action/replication_record"}},"created_at":"2026-07-05T11:08:47.965181+00:00","updated_at":"2026-07-05T11:08:47.965181+00:00"}