{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:HFTS3HNIMVCDCC4VQEZ7C7MTR5","short_pith_number":"pith:HFTS3HNI","schema_version":"1.0","canonical_sha256":"39672d9da86544310b958133f17d938f59be62cb2b71b8e93e4d3b040cee2da1","source":{"kind":"arxiv","id":"2412.09596","version":1},"attestation_state":"computed","paper":{"title":"InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.CV","authors_text":"Bin Wang, Conghui He, Dahua Lin, Han Lv, Haodong Duan, Jiaqi Wang, Jiaye Ge, Jingwen Li, Junbo Niu, Kai Chen, Lin Chen, Min Zhang, Pan Zhang, Qipeng Guo, Rui Qian, Shuangrui Ding, Wei Li, Wenwei Zhang, Xiaoyi Dong, Xilin Wei, Xin Chen, Xingcheng Zhang, Xinyue Zhang, Yifei Li, Yuhang Cao, Yuhang Zang, Yu Qiao, Zheng Nie, Zhongying Tu","submitted_at":"2024-12-12T18:58:30Z","abstract_excerpt":"Creating AI systems that can interact with environments over long periods, similar to human cognition, has been a longstanding research goal. Recent advancements in multimodal large language models (MLLMs) have made significant strides in open-world understanding. However, the challenge of continuous and simultaneous streaming perception, memory, and reasoning remains largely unexplored. Current MLLMs are constrained by their sequence-to-sequence architecture, which limits their ability to process inputs and generate responses simultaneously, akin to being unable to think while perceiving. Fur"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2412.09596","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2024-12-12T18:58:30Z","cross_cats_sorted":["cs.AI","cs.CL"],"title_canon_sha256":"dca6f56847148cc0dcf662908f03e5499284fbe7dfff47db6c78a25519c565e0","abstract_canon_sha256":"da671c6065521edee2c1ebc5fa797ef678ce7330bfc76a19b0f718d429361465"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:48:26.869640Z","signature_b64":"ll82BJ6kxvFF7XsHrxgAsOb8iPE4zaiGQZXqGtNPqnvdU0dV2kgXHmrhEh2gz8MQ7+4MMEepgx2fgy4XW1HTBQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"39672d9da86544310b958133f17d938f59be62cb2b71b8e93e4d3b040cee2da1","last_reissued_at":"2026-07-05T09:48:26.869146Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:48:26.869146Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.CL"],"primary_cat":"cs.CV","authors_text":"Bin Wang, Conghui He, Dahua Lin, Han Lv, Haodong Duan, Jiaqi Wang, Jiaye Ge, Jingwen Li, Junbo Niu, Kai Chen, Lin Chen, Min Zhang, Pan Zhang, Qipeng Guo, Rui Qian, Shuangrui Ding, Wei Li, Wenwei Zhang, Xiaoyi Dong, Xilin Wei, Xin Chen, Xingcheng Zhang, Xinyue Zhang, Yifei Li, Yuhang Cao, Yuhang Zang, Yu Qiao, Zheng Nie, Zhongying Tu","submitted_at":"2024-12-12T18:58:30Z","abstract_excerpt":"Creating AI systems that can interact with environments over long periods, similar to human cognition, has been a longstanding research goal. Recent advancements in multimodal large language models (MLLMs) have made significant strides in open-world understanding. However, the challenge of continuous and simultaneous streaming perception, memory, and reasoning remains largely unexplored. Current MLLMs are constrained by their sequence-to-sequence architecture, which limits their ability to process inputs and generate responses simultaneously, akin to being unable to think while perceiving. Fur"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2412.09596","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2412.09596/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2412.09596","created_at":"2026-07-05T09:48:26.869204+00:00"},{"alias_kind":"arxiv_version","alias_value":"2412.09596v1","created_at":"2026-07-05T09:48:26.869204+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2412.09596","created_at":"2026-07-05T09:48:26.869204+00:00"},{"alias_kind":"pith_short_12","alias_value":"HFTS3HNIMVCD","created_at":"2026-07-05T09:48:26.869204+00:00"},{"alias_kind":"pith_short_16","alias_value":"HFTS3HNIMVCDCC4V","created_at":"2026-07-05T09:48:26.869204+00:00"},{"alias_kind":"pith_short_8","alias_value":"HFTS3HNI","created_at":"2026-07-05T09:48:26.869204+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.19849","citing_title":"ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2606.19338","citing_title":"Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games","ref_index":92,"is_internal_anchor":false},{"citing_arxiv_id":"2512.15693","citing_title":"Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning","ref_index":84,"is_internal_anchor":false},{"citing_arxiv_id":"2503.01743","citing_title":"Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs","ref_index":55,"is_internal_anchor":false},{"citing_arxiv_id":"2604.13804","citing_title":"Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning","ref_index":44,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/HFTS3HNIMVCDCC4VQEZ7C7MTR5","json":"https://pith.science/pith/HFTS3HNIMVCDCC4VQEZ7C7MTR5.json","graph_json":"https://pith.science/api/pith-number/HFTS3HNIMVCDCC4VQEZ7C7MTR5/graph.json","events_json":"https://pith.science/api/pith-number/HFTS3HNIMVCDCC4VQEZ7C7MTR5/events.json","paper":"https://pith.science/paper/HFTS3HNI"},"agent_actions":{"view_html":"https://pith.science/pith/HFTS3HNIMVCDCC4VQEZ7C7MTR5","download_json":"https://pith.science/pith/HFTS3HNIMVCDCC4VQEZ7C7MTR5.json","view_paper":"https://pith.science/paper/HFTS3HNI","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2412.09596&json=true","fetch_graph":"https://pith.science/api/pith-number/HFTS3HNIMVCDCC4VQEZ7C7MTR5/graph.json","fetch_events":"https://pith.science/api/pith-number/HFTS3HNIMVCDCC4VQEZ7C7MTR5/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/HFTS3HNIMVCDCC4VQEZ7C7MTR5/action/timestamp_anchor","attest_storage":"https://pith.science/pith/HFTS3HNIMVCDCC4VQEZ7C7MTR5/action/storage_attestation","attest_author":"https://pith.science/pith/HFTS3HNIMVCDCC4VQEZ7C7MTR5/action/author_attestation","sign_citation":"https://pith.science/pith/HFTS3HNIMVCDCC4VQEZ7C7MTR5/action/citation_signature","submit_replication":"https://pith.science/pith/HFTS3HNIMVCDCC4VQEZ7C7MTR5/action/replication_record"}},"created_at":"2026-07-05T09:48:26.869204+00:00","updated_at":"2026-07-05T09:48:26.869204+00:00"}