{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:KTAPXYFCU77XLKAEEHBQ2Z3WLZ","short_pith_number":"pith:KTAPXYFC","schema_version":"1.0","canonical_sha256":"54c0fbe0a2a7ff75a80421c30d67765e6a74422d183460ce01bb54a76203707a","source":{"kind":"arxiv","id":"2508.08601","version":3},"attestation_state":"computed","paper":{"title":"Yan: Foundational Interactive Video Generation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Deheng Ye, Fangyun Zhou, Jiacheng Lv, Jianqi Ma, Junyan Lv, Junyou Li, Jun Zhang, Mingyu Yang, Minwen Deng, Qiang Fu, Wei Yang, Wenkai Lv, Yangbin Yu, Yewen Wang, Yonghang Guan, Zhihao Hu, Zhongbin Fang, Zhongqian Sun","submitted_at":"2025-08-12T03:34:21Z","abstract_excerpt":"We present Yan, a foundational framework for interactive video generation, covering the entire pipeline from simulation and generation to editing. Specifically, Yan comprises three core modules. AAA-level Simulation: We design a highly-compressed, low-latency 3D-VAE coupled with a KV-cache-based shift-window denoising inference process, achieving real-time 1080P/60FPS interactive simulation. Multi-Modal Generation: We introduce a hierarchical autoregressive caption method that injects game-specific knowledge into open-domain multi-modal video diffusion models (VDMs), then transforming the VDM "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2508.08601","kind":"arxiv","version":3},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CV","submitted_at":"2025-08-12T03:34:21Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"b514c4b046f0eb8dff076331d06e000e0b4aa1bc3befc293c942d1ac8e44226f","abstract_canon_sha256":"004703667737dfde97c1fdb625d42740c97dc7c17f27ef942d8623506016135f"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:53:50.609636Z","signature_b64":"LJKwPcaF261MaG3x2JglHxkoSQEG4q2MU7LAf+w7HIfU+/0UAjObFIReJt4U41hE8vjIgg6OGVZzR7fQkZ/8DQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"54c0fbe0a2a7ff75a80421c30d67765e6a74422d183460ce01bb54a76203707a","last_reissued_at":"2026-07-05T11:53:50.609139Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:53:50.609139Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Yan: Foundational Interactive Video Generation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CV","authors_text":"Deheng Ye, Fangyun Zhou, Jiacheng Lv, Jianqi Ma, Junyan Lv, Junyou Li, Jun Zhang, Mingyu Yang, Minwen Deng, Qiang Fu, Wei Yang, Wenkai Lv, Yangbin Yu, Yewen Wang, Yonghang Guan, Zhihao Hu, Zhongbin Fang, Zhongqian Sun","submitted_at":"2025-08-12T03:34:21Z","abstract_excerpt":"We present Yan, a foundational framework for interactive video generation, covering the entire pipeline from simulation and generation to editing. Specifically, Yan comprises three core modules. AAA-level Simulation: We design a highly-compressed, low-latency 3D-VAE coupled with a KV-cache-based shift-window denoising inference process, achieving real-time 1080P/60FPS interactive simulation. Multi-Modal Generation: We introduce a hierarchical autoregressive caption method that injects game-specific knowledge into open-domain multi-modal video diffusion models (VDMs), then transforming the VDM "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2508.08601","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2508.08601/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2508.08601","created_at":"2026-07-05T11:53:50.609199+00:00"},{"alias_kind":"arxiv_version","alias_value":"2508.08601v3","created_at":"2026-07-05T11:53:50.609199+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2508.08601","created_at":"2026-07-05T11:53:50.609199+00:00"},{"alias_kind":"pith_short_12","alias_value":"KTAPXYFCU77X","created_at":"2026-07-05T11:53:50.609199+00:00"},{"alias_kind":"pith_short_16","alias_value":"KTAPXYFCU77XLKAE","created_at":"2026-07-05T11:53:50.609199+00:00"},{"alias_kind":"pith_short_8","alias_value":"KTAPXYFC","created_at":"2026-07-05T11:53:50.609199+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":11,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.07967","citing_title":"DisCo: World Models with Discrete Camera Motion Control","ref_index":48,"is_internal_anchor":false},{"citing_arxiv_id":"2606.07508","citing_title":"Streaming Video Generation with Streaming Force Control","ref_index":77,"is_internal_anchor":false},{"citing_arxiv_id":"2606.07326","citing_title":"AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization","ref_index":54,"is_internal_anchor":false},{"citing_arxiv_id":"2606.01164","citing_title":"Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends","ref_index":204,"is_internal_anchor":false},{"citing_arxiv_id":"2605.30263","citing_title":"minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2605.31336","citing_title":"DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory","ref_index":47,"is_internal_anchor":false},{"citing_arxiv_id":"2605.22718","citing_title":"WorldKV: Efficient World Memory with World Retrieval and Compression","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2512.21714","citing_title":"AstraNav-World: World Model for Foresight Control and Consistency","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2602.06949","citing_title":"DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos","ref_index":113,"is_internal_anchor":false},{"citing_arxiv_id":"2509.24527","citing_title":"Training Agents Inside of Scalable World Models","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2604.15911","citing_title":"Efficient Video Diffusion Models: Advancements and Challenges","ref_index":160,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/KTAPXYFCU77XLKAEEHBQ2Z3WLZ","json":"https://pith.science/pith/KTAPXYFCU77XLKAEEHBQ2Z3WLZ.json","graph_json":"https://pith.science/api/pith-number/KTAPXYFCU77XLKAEEHBQ2Z3WLZ/graph.json","events_json":"https://pith.science/api/pith-number/KTAPXYFCU77XLKAEEHBQ2Z3WLZ/events.json","paper":"https://pith.science/paper/KTAPXYFC"},"agent_actions":{"view_html":"https://pith.science/pith/KTAPXYFCU77XLKAEEHBQ2Z3WLZ","download_json":"https://pith.science/pith/KTAPXYFCU77XLKAEEHBQ2Z3WLZ.json","view_paper":"https://pith.science/paper/KTAPXYFC","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2508.08601&json=true","fetch_graph":"https://pith.science/api/pith-number/KTAPXYFCU77XLKAEEHBQ2Z3WLZ/graph.json","fetch_events":"https://pith.science/api/pith-number/KTAPXYFCU77XLKAEEHBQ2Z3WLZ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/KTAPXYFCU77XLKAEEHBQ2Z3WLZ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/KTAPXYFCU77XLKAEEHBQ2Z3WLZ/action/storage_attestation","attest_author":"https://pith.science/pith/KTAPXYFCU77XLKAEEHBQ2Z3WLZ/action/author_attestation","sign_citation":"https://pith.science/pith/KTAPXYFCU77XLKAEEHBQ2Z3WLZ/action/citation_signature","submit_replication":"https://pith.science/pith/KTAPXYFCU77XLKAEEHBQ2Z3WLZ/action/replication_record"}},"created_at":"2026-07-05T11:53:50.609199+00:00","updated_at":"2026-07-05T11:53:50.609199+00:00"}