{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:F5GMLN42OAAW6ZLVWLCFZDLBMU","short_pith_number":"pith:F5GMLN42","schema_version":"1.0","canonical_sha256":"2f4cc5b79a70016f6575b2c45c8d61651f50351f1d4c3243d31837c4363a0977","source":{"kind":"arxiv","id":"2504.21277","version":2},"attestation_state":"computed","paper":{"title":"Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.AI","authors_text":"Cen Chen, Guanghao Zhou, Jian Xu, Jie Wang, Minghui Qiu, Panjia Qiu, Zheming Yang","submitted_at":"2025-04-30T03:14:28Z","abstract_excerpt":"The application of reinforcement learning (RL) to enhance the reasoning capabilities of Multimodal Large Language Models (MLLMs) constitutes a rapidly advancing research area. While MLLMs extend Large Language Models (LLMs) to handle diverse modalities such as vision, audio, and video, enabling robust reasoning across multimodal inputs remains challenging. This paper provides a systematic review of recent advances in RL-based reasoning for MLLMs, covering key algorithmic designs, reward mechanism innovations, and practical applications. We highlight two main RL paradigms, value-model-free and "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2504.21277","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.AI","submitted_at":"2025-04-30T03:14:28Z","cross_cats_sorted":[],"title_canon_sha256":"1195d9b47d92859b8be5b965a478e076f01123c836a052a18d884d38822964c5","abstract_canon_sha256":"4575f098458dba8306ad39506e15d0296667e8f8dd5c2078ac1f8887301a7526"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:06:31.219206Z","signature_b64":"JZz+g6JJ0vrXggWUHX/OXjZ1VPo/PpdR43XkuBc/rixF1EhuThYGORWBjSvzqwrqTEbngwj2YxIy9/ZLjT3HDg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"2f4cc5b79a70016f6575b2c45c8d61651f50351f1d4c3243d31837c4363a0977","last_reissued_at":"2026-07-05T11:06:31.218704Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:06:31.218704Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.AI","authors_text":"Cen Chen, Guanghao Zhou, Jian Xu, Jie Wang, Minghui Qiu, Panjia Qiu, Zheming Yang","submitted_at":"2025-04-30T03:14:28Z","abstract_excerpt":"The application of reinforcement learning (RL) to enhance the reasoning capabilities of Multimodal Large Language Models (MLLMs) constitutes a rapidly advancing research area. While MLLMs extend Large Language Models (LLMs) to handle diverse modalities such as vision, audio, and video, enabling robust reasoning across multimodal inputs remains challenging. This paper provides a systematic review of recent advances in RL-based reasoning for MLLMs, covering key algorithmic designs, reward mechanism innovations, and practical applications. We highlight two main RL paradigms, value-model-free and "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2504.21277","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2504.21277/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2504.21277","created_at":"2026-07-05T11:06:31.218764+00:00"},{"alias_kind":"arxiv_version","alias_value":"2504.21277v2","created_at":"2026-07-05T11:06:31.218764+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2504.21277","created_at":"2026-07-05T11:06:31.218764+00:00"},{"alias_kind":"pith_short_12","alias_value":"F5GMLN42OAAW","created_at":"2026-07-05T11:06:31.218764+00:00"},{"alias_kind":"pith_short_16","alias_value":"F5GMLN42OAAW6ZLV","created_at":"2026-07-05T11:06:31.218764+00:00"},{"alias_kind":"pith_short_8","alias_value":"F5GMLN42","created_at":"2026-07-05T11:06:31.218764+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":14,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.08572","citing_title":"Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning","ref_index":23,"is_internal_anchor":true},{"citing_arxiv_id":"2606.18988","citing_title":"ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection","ref_index":46,"is_internal_anchor":false},{"citing_arxiv_id":"2606.18216","citing_title":"Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients","ref_index":85,"is_internal_anchor":false},{"citing_arxiv_id":"2605.25571","citing_title":"AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2605.15951","citing_title":"From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding","ref_index":102,"is_internal_anchor":false},{"citing_arxiv_id":"2604.27859","citing_title":"Rethinking Agentic Reinforcement Learning In Large Language Models","ref_index":130,"is_internal_anchor":false},{"citing_arxiv_id":"2509.02547","citing_title":"The Landscape of Agentic Reinforcement Learning for LLMs: A Survey","ref_index":223,"is_internal_anchor":false},{"citing_arxiv_id":"2512.03043","citing_title":"OneThinker: All-in-one Reasoning Model for Image and Video","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2602.00181","citing_title":"CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning","ref_index":51,"is_internal_anchor":false},{"citing_arxiv_id":"2604.27859","citing_title":"Rethinking Agentic Reinforcement Learning In Large Language Models","ref_index":130,"is_internal_anchor":false},{"citing_arxiv_id":"2604.27859","citing_title":"Rethinking Agentic Reinforcement Learning In Large Language Models","ref_index":130,"is_internal_anchor":false},{"citing_arxiv_id":"2604.19264","citing_title":"DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2604.08209","citing_title":"OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering","ref_index":57,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07872","citing_title":"Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models","ref_index":52,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/F5GMLN42OAAW6ZLVWLCFZDLBMU","json":"https://pith.science/pith/F5GMLN42OAAW6ZLVWLCFZDLBMU.json","graph_json":"https://pith.science/api/pith-number/F5GMLN42OAAW6ZLVWLCFZDLBMU/graph.json","events_json":"https://pith.science/api/pith-number/F5GMLN42OAAW6ZLVWLCFZDLBMU/events.json","paper":"https://pith.science/paper/F5GMLN42"},"agent_actions":{"view_html":"https://pith.science/pith/F5GMLN42OAAW6ZLVWLCFZDLBMU","download_json":"https://pith.science/pith/F5GMLN42OAAW6ZLVWLCFZDLBMU.json","view_paper":"https://pith.science/paper/F5GMLN42","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2504.21277&json=true","fetch_graph":"https://pith.science/api/pith-number/F5GMLN42OAAW6ZLVWLCFZDLBMU/graph.json","fetch_events":"https://pith.science/api/pith-number/F5GMLN42OAAW6ZLVWLCFZDLBMU/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/F5GMLN42OAAW6ZLVWLCFZDLBMU/action/timestamp_anchor","attest_storage":"https://pith.science/pith/F5GMLN42OAAW6ZLVWLCFZDLBMU/action/storage_attestation","attest_author":"https://pith.science/pith/F5GMLN42OAAW6ZLVWLCFZDLBMU/action/author_attestation","sign_citation":"https://pith.science/pith/F5GMLN42OAAW6ZLVWLCFZDLBMU/action/citation_signature","submit_replication":"https://pith.science/pith/F5GMLN42OAAW6ZLVWLCFZDLBMU/action/replication_record"}},"created_at":"2026-07-05T11:06:31.218764+00:00","updated_at":"2026-07-05T11:06:31.218764+00:00"}