{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:5CYMVCK4QPO5S7CTNHTQ4OCHEJ","short_pith_number":"pith:5CYMVCK4","schema_version":"1.0","canonical_sha256":"e8b0ca895c83ddd97c5369e70e3847224a642f6a46abc3d6469029e14c4ecd0e","source":{"kind":"arxiv","id":"2504.07954","version":1},"attestation_state":"computed","paper":{"title":"Perception-R1: Pioneering Perception Policy with Reinforcement Learning","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.CV","authors_text":"Chunrui Han, Daxin Jiang, En Yu, Haoran Wei, Jianjian Sun, Jingyu Wang, Jisheng Yin, Kangheng Lin, Liang Zhao, Wenbing Tao, Xiangyu Zhang, Yana Wei, Yuang Peng, Zheng Ge","submitted_at":"2025-04-10T17:58:27Z","abstract_excerpt":"Inspired by the success of DeepSeek-R1, we explore the potential of rule-based reinforcement learning (RL) in MLLM post-training for perception policy learning. While promising, our initial experiments reveal that incorporating a thinking process through RL does not consistently lead to performance gains across all visual perception tasks. This leads us to delve into the essential role of RL in the context of visual perception. In this work, we return to the fundamentals and explore the effects of RL on different perception tasks. We observe that the perceptual complexity is a major factor in "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2504.07954","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2025-04-10T17:58:27Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"ec28f7582c01353b9a9f63ff690d5bb183a00b077c7671eca485145ec66dcea5","abstract_canon_sha256":"ea8d4edd31fc0a4d1478f9748dfb91985b853485d7804b9e3bb5d86c33e1b68f"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:47:22.558512Z","signature_b64":"9uDKFvBVPWV7eQprcKCs2YtiPZGR0AQGNiuKBKvPmbeq4Uwa3+2/u3lhoFom3FaXDbwuu01nxLWzInKC10YwCw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"e8b0ca895c83ddd97c5369e70e3847224a642f6a46abc3d6469029e14c4ecd0e","last_reissued_at":"2026-07-05T10:47:22.558027Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:47:22.558027Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Perception-R1: Pioneering Perception Policy with Reinforcement Learning","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.CV","authors_text":"Chunrui Han, Daxin Jiang, En Yu, Haoran Wei, Jianjian Sun, Jingyu Wang, Jisheng Yin, Kangheng Lin, Liang Zhao, Wenbing Tao, Xiangyu Zhang, Yana Wei, Yuang Peng, Zheng Ge","submitted_at":"2025-04-10T17:58:27Z","abstract_excerpt":"Inspired by the success of DeepSeek-R1, we explore the potential of rule-based reinforcement learning (RL) in MLLM post-training for perception policy learning. While promising, our initial experiments reveal that incorporating a thinking process through RL does not consistently lead to performance gains across all visual perception tasks. This leads us to delve into the essential role of RL in the context of visual perception. In this work, we return to the fundamentals and explore the effects of RL on different perception tasks. We observe that the perceptual complexity is a major factor in "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2504.07954","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2504.07954/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2504.07954","created_at":"2026-07-05T10:47:22.558085+00:00"},{"alias_kind":"arxiv_version","alias_value":"2504.07954v1","created_at":"2026-07-05T10:47:22.558085+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2504.07954","created_at":"2026-07-05T10:47:22.558085+00:00"},{"alias_kind":"pith_short_12","alias_value":"5CYMVCK4QPO5","created_at":"2026-07-05T10:47:22.558085+00:00"},{"alias_kind":"pith_short_16","alias_value":"5CYMVCK4QPO5S7CT","created_at":"2026-07-05T10:47:22.558085+00:00"},{"alias_kind":"pith_short_8","alias_value":"5CYMVCK4","created_at":"2026-07-05T10:47:22.558085+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":24,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2604.18483","citing_title":"Steadily moving semi-infinite fracture in plane poroelasticity","ref_index":110,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26196","citing_title":"From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models","ref_index":171,"is_internal_anchor":false},{"citing_arxiv_id":"2607.01191","citing_title":"Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2605.29198","citing_title":"Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization","ref_index":46,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14068","citing_title":"CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2605.16079","citing_title":"VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocation","ref_index":44,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18903","citing_title":"Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era","ref_index":58,"is_internal_anchor":false},{"citing_arxiv_id":"2605.20165","citing_title":"CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models","ref_index":51,"is_internal_anchor":false},{"citing_arxiv_id":"2507.00748","citing_title":"Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning","ref_index":47,"is_internal_anchor":false},{"citing_arxiv_id":"2511.11113","citing_title":"VIDEOP2R: Video Understanding from Perception to Reasoning","ref_index":64,"is_internal_anchor":false},{"citing_arxiv_id":"2512.03043","citing_title":"OneThinker: All-in-one Reasoning Model for Image and Video","ref_index":14,"is_internal_anchor":false},{"citing_arxiv_id":"2512.06721","citing_title":"ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems in the Wild","ref_index":61,"is_internal_anchor":false},{"citing_arxiv_id":"2601.06993","citing_title":"Can Textual Reasoning Improve the Performance of MLLMs on Fine-grained Visual Classification?","ref_index":43,"is_internal_anchor":false},{"citing_arxiv_id":"2602.00181","citing_title":"CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning","ref_index":49,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14068","citing_title":"CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves","ref_index":42,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12163","citing_title":"Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model","ref_index":56,"is_internal_anchor":false},{"citing_arxiv_id":"2604.03307","citing_title":"V-Reflection: Transforming MLLMs from Passive Observers to Active Interrogators","ref_index":35,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12163","citing_title":"Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model","ref_index":56,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01402","citing_title":"Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2604.22498","citing_title":"CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding","ref_index":54,"is_internal_anchor":false},{"citing_arxiv_id":"2605.01402","citing_title":"Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2604.20705","citing_title":"SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models","ref_index":78,"is_internal_anchor":false},{"citing_arxiv_id":"2604.08545","citing_title":"Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models","ref_index":47,"is_internal_anchor":false},{"citing_arxiv_id":"2604.18484","citing_title":"XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments","ref_index":110,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/5CYMVCK4QPO5S7CTNHTQ4OCHEJ","json":"https://pith.science/pith/5CYMVCK4QPO5S7CTNHTQ4OCHEJ.json","graph_json":"https://pith.science/api/pith-number/5CYMVCK4QPO5S7CTNHTQ4OCHEJ/graph.json","events_json":"https://pith.science/api/pith-number/5CYMVCK4QPO5S7CTNHTQ4OCHEJ/events.json","paper":"https://pith.science/paper/5CYMVCK4"},"agent_actions":{"view_html":"https://pith.science/pith/5CYMVCK4QPO5S7CTNHTQ4OCHEJ","download_json":"https://pith.science/pith/5CYMVCK4QPO5S7CTNHTQ4OCHEJ.json","view_paper":"https://pith.science/paper/5CYMVCK4","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2504.07954&json=true","fetch_graph":"https://pith.science/api/pith-number/5CYMVCK4QPO5S7CTNHTQ4OCHEJ/graph.json","fetch_events":"https://pith.science/api/pith-number/5CYMVCK4QPO5S7CTNHTQ4OCHEJ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/5CYMVCK4QPO5S7CTNHTQ4OCHEJ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/5CYMVCK4QPO5S7CTNHTQ4OCHEJ/action/storage_attestation","attest_author":"https://pith.science/pith/5CYMVCK4QPO5S7CTNHTQ4OCHEJ/action/author_attestation","sign_citation":"https://pith.science/pith/5CYMVCK4QPO5S7CTNHTQ4OCHEJ/action/citation_signature","submit_replication":"https://pith.science/pith/5CYMVCK4QPO5S7CTNHTQ4OCHEJ/action/replication_record"}},"created_at":"2026-07-05T10:47:22.558085+00:00","updated_at":"2026-07-05T10:47:22.558085+00:00"}