{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:27WCFWZGSZS6HM7WGEHYD6US7K","short_pith_number":"pith:27WCFWZG","schema_version":"1.0","canonical_sha256":"d7ec22db269665e3b3f6310f81fa92fa9a9ef808f6f70e1c9554c898035eff49","source":{"kind":"arxiv","id":"2506.21277","version":1},"attestation_state":"computed","paper":{"title":"HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context","license":"http://creativecommons.org/licenses/by-nc-nd/4.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.CV","authors_text":"Bowen Yin, Boyuan Sun, Detao Bai, Jiaxing Zhao, Jingren Zhou, Qize Yang, Shenghao Fu, Shimin Yao, Weixuan Chen, Xihan Wei","submitted_at":"2025-06-26T14:01:03Z","abstract_excerpt":"With the rapid evolution of multimodal large language models, the capacity to deeply understand and interpret human intentions has emerged as a critical capability, which demands detailed and thoughtful reasoning. In recent studies, Reinforcement Learning (RL) has demonstrated potential in enhancing the reasoning capabilities of Large Language Models (LLMs). Nonetheless, the challenges associated with adapting RL to multimodal data and formats remain largely unaddressed. In this paper, we identify two issues in existing multimodal reasoning models: insufficient global context understanding and"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2506.21277","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-nd/4.0/","primary_cat":"cs.CV","submitted_at":"2025-06-26T14:01:03Z","cross_cats_sorted":["cs.CL"],"title_canon_sha256":"3216e2bed018f29d3bec0af9bf22372ea39c67030578f8626baf62200cdfeb7a","abstract_canon_sha256":"e9a36e70154b4c07da49fffa2bda922d2a85da05a9f44f357461382d847d51d4"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:27:42.778700Z","signature_b64":"nfYIVRHSsp6O00dnj3pRzaAmRaIX52NULZD2FVXPgC5k1i8s9wvFK6o6+FwhUWpAxC4X8iWsE6Ul4MlsDtreAA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"d7ec22db269665e3b3f6310f81fa92fa9a9ef808f6f70e1c9554c898035eff49","last_reissued_at":"2026-07-05T11:27:42.778149Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:27:42.778149Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context","license":"http://creativecommons.org/licenses/by-nc-nd/4.0/","headline":"","cross_cats":["cs.CL"],"primary_cat":"cs.CV","authors_text":"Bowen Yin, Boyuan Sun, Detao Bai, Jiaxing Zhao, Jingren Zhou, Qize Yang, Shenghao Fu, Shimin Yao, Weixuan Chen, Xihan Wei","submitted_at":"2025-06-26T14:01:03Z","abstract_excerpt":"With the rapid evolution of multimodal large language models, the capacity to deeply understand and interpret human intentions has emerged as a critical capability, which demands detailed and thoughtful reasoning. In recent studies, Reinforcement Learning (RL) has demonstrated potential in enhancing the reasoning capabilities of Large Language Models (LLMs). Nonetheless, the challenges associated with adapting RL to multimodal data and formats remain largely unaddressed. In this paper, we identify two issues in existing multimodal reasoning models: insufficient global context understanding and"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2506.21277","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2506.21277/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2506.21277","created_at":"2026-07-05T11:27:42.778210+00:00"},{"alias_kind":"arxiv_version","alias_value":"2506.21277v1","created_at":"2026-07-05T11:27:42.778210+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2506.21277","created_at":"2026-07-05T11:27:42.778210+00:00"},{"alias_kind":"pith_short_12","alias_value":"27WCFWZGSZS6","created_at":"2026-07-05T11:27:42.778210+00:00"},{"alias_kind":"pith_short_16","alias_value":"27WCFWZGSZS6HM7W","created_at":"2026-07-05T11:27:42.778210+00:00"},{"alias_kind":"pith_short_8","alias_value":"27WCFWZG","created_at":"2026-07-05T11:27:42.778210+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":16,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.25325","citing_title":"Omni-Perception Policy Optimization for Multimodal Emotion Reasoning","ref_index":21,"is_internal_anchor":false},{"citing_arxiv_id":"2606.20970","citing_title":"CogniRoute: Learning to Route Social Evidence in Omni-Modal Models","ref_index":91,"is_internal_anchor":false},{"citing_arxiv_id":"2607.01667","citing_title":"Temporal and Cross-Modal Alignment for Enhanced Audiovisual Video Captioning","ref_index":64,"is_internal_anchor":false},{"citing_arxiv_id":"2606.12018","citing_title":"MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2606.27652","citing_title":"MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22012","citing_title":"LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning","ref_index":55,"is_internal_anchor":false},{"citing_arxiv_id":"2605.18018","citing_title":"See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding","ref_index":81,"is_internal_anchor":false},{"citing_arxiv_id":"2511.14582","citing_title":"OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models","ref_index":58,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12034","citing_title":"Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12056","citing_title":"OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models","ref_index":55,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12034","citing_title":"Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2605.09996","citing_title":"Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2604.11244","citing_title":"Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding","ref_index":30,"is_internal_anchor":false},{"citing_arxiv_id":"2604.08209","citing_title":"OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering","ref_index":47,"is_internal_anchor":false},{"citing_arxiv_id":"2604.14520","citing_title":"Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs","ref_index":33,"is_internal_anchor":false},{"citing_arxiv_id":"2604.16617","citing_title":"AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers","ref_index":36,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/27WCFWZGSZS6HM7WGEHYD6US7K","json":"https://pith.science/pith/27WCFWZGSZS6HM7WGEHYD6US7K.json","graph_json":"https://pith.science/api/pith-number/27WCFWZGSZS6HM7WGEHYD6US7K/graph.json","events_json":"https://pith.science/api/pith-number/27WCFWZGSZS6HM7WGEHYD6US7K/events.json","paper":"https://pith.science/paper/27WCFWZG"},"agent_actions":{"view_html":"https://pith.science/pith/27WCFWZGSZS6HM7WGEHYD6US7K","download_json":"https://pith.science/pith/27WCFWZGSZS6HM7WGEHYD6US7K.json","view_paper":"https://pith.science/paper/27WCFWZG","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2506.21277&json=true","fetch_graph":"https://pith.science/api/pith-number/27WCFWZGSZS6HM7WGEHYD6US7K/graph.json","fetch_events":"https://pith.science/api/pith-number/27WCFWZGSZS6HM7WGEHYD6US7K/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/27WCFWZGSZS6HM7WGEHYD6US7K/action/timestamp_anchor","attest_storage":"https://pith.science/pith/27WCFWZGSZS6HM7WGEHYD6US7K/action/storage_attestation","attest_author":"https://pith.science/pith/27WCFWZGSZS6HM7WGEHYD6US7K/action/author_attestation","sign_citation":"https://pith.science/pith/27WCFWZGSZS6HM7WGEHYD6US7K/action/citation_signature","submit_replication":"https://pith.science/pith/27WCFWZGSZS6HM7WGEHYD6US7K/action/replication_record"}},"created_at":"2026-07-05T11:27:42.778210+00:00","updated_at":"2026-07-05T11:27:42.778210+00:00"}