{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:AIPTQ63NI2WFYE33MHJR6KN66O","short_pith_number":"pith:AIPTQ63N","schema_version":"1.0","canonical_sha256":"021f387b6d46ac5c137b61d31f29bef3826ab44f7851dd174f66e2d9b880e4cc","source":{"kind":"arxiv","id":"2410.23277","version":2},"attestation_state":"computed","paper":{"title":"SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CL","cs.LG","cs.RO"],"primary_cat":"cs.CV","authors_text":"Beide Liu, Chung-Ching Lin, Jianfeng Wang, Kai-Wei Chang, Kevin Lin, Lijuan Wang, Linjie Li, Maxine Wu, Yingnian Wu, Yining Hong, Yuanhao Zhai, Zhengyuan Yang","submitted_at":"2024-10-30T17:55:52Z","abstract_excerpt":"Human beings are endowed with a complementary learning system, which bridges the slow learning of general world dynamics with fast storage of episodic memory from a new experience. Previous video generation models, however, primarily focus on slow learning by pre-training on vast amounts of data, overlooking the fast learning phase crucial for episodic memory storage. This oversight leads to inconsistencies across temporally distant frames when generating longer videos, as these frames fall beyond the model's context window. To this end, we introduce SlowFast-VGen, a novel dual-speed learning "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2410.23277","kind":"arxiv","version":2},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CV","submitted_at":"2024-10-30T17:55:52Z","cross_cats_sorted":["cs.AI","cs.CL","cs.LG","cs.RO"],"title_canon_sha256":"cd5dbd899bded839155ea83fe585648f5b9c0ae96a5d5b3439b043eb5e62b3c6","abstract_canon_sha256":"5c28f8a0e63db0fae19545e60d7dbfc26365c7da1e226b2f95c7278e2abeb024"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T09:29:35.501070Z","signature_b64":"MJ+v0+HOreV6vRrhOfXRn7H1/6cOLOh2YARK2i7kNIeKrNBUVMK9GY7ul+SZiHQpSvCzYLXb1IFgk3J5U7IaBw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"021f387b6d46ac5c137b61d31f29bef3826ab44f7851dd174f66e2d9b880e4cc","last_reissued_at":"2026-07-05T09:29:35.500626Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T09:29:35.500626Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI","cs.CL","cs.LG","cs.RO"],"primary_cat":"cs.CV","authors_text":"Beide Liu, Chung-Ching Lin, Jianfeng Wang, Kai-Wei Chang, Kevin Lin, Lijuan Wang, Linjie Li, Maxine Wu, Yingnian Wu, Yining Hong, Yuanhao Zhai, Zhengyuan Yang","submitted_at":"2024-10-30T17:55:52Z","abstract_excerpt":"Human beings are endowed with a complementary learning system, which bridges the slow learning of general world dynamics with fast storage of episodic memory from a new experience. Previous video generation models, however, primarily focus on slow learning by pre-training on vast amounts of data, overlooking the fast learning phase crucial for episodic memory storage. This oversight leads to inconsistencies across temporally distant frames when generating longer videos, as these frames fall beyond the model's context window. To this end, we introduce SlowFast-VGen, a novel dual-speed learning "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2410.23277","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2410.23277/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2410.23277","created_at":"2026-07-05T09:29:35.500678+00:00"},{"alias_kind":"arxiv_version","alias_value":"2410.23277v2","created_at":"2026-07-05T09:29:35.500678+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2410.23277","created_at":"2026-07-05T09:29:35.500678+00:00"},{"alias_kind":"pith_short_12","alias_value":"AIPTQ63NI2WF","created_at":"2026-07-05T09:29:35.500678+00:00"},{"alias_kind":"pith_short_16","alias_value":"AIPTQ63NI2WFYE33","created_at":"2026-07-05T09:29:35.500678+00:00"},{"alias_kind":"pith_short_8","alias_value":"AIPTQ63N","created_at":"2026-07-05T09:29:35.500678+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":5,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.22370","citing_title":"Towards Error-Free Long Video Generation","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2505.21996","citing_title":"VRAG: Learning World Models for Interactive Video Generation","ref_index":34,"is_internal_anchor":false},{"citing_arxiv_id":"2512.15840","citing_title":"Large Video Planner Enables Generalizable Robot Control","ref_index":39,"is_internal_anchor":false},{"citing_arxiv_id":"2605.12496","citing_title":"CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives","ref_index":15,"is_internal_anchor":false},{"citing_arxiv_id":"2604.13036","citing_title":"Lyra 2.0: Explorable Generative 3D Worlds","ref_index":32,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/AIPTQ63NI2WFYE33MHJR6KN66O","json":"https://pith.science/pith/AIPTQ63NI2WFYE33MHJR6KN66O.json","graph_json":"https://pith.science/api/pith-number/AIPTQ63NI2WFYE33MHJR6KN66O/graph.json","events_json":"https://pith.science/api/pith-number/AIPTQ63NI2WFYE33MHJR6KN66O/events.json","paper":"https://pith.science/paper/AIPTQ63N"},"agent_actions":{"view_html":"https://pith.science/pith/AIPTQ63NI2WFYE33MHJR6KN66O","download_json":"https://pith.science/pith/AIPTQ63NI2WFYE33MHJR6KN66O.json","view_paper":"https://pith.science/paper/AIPTQ63N","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2410.23277&json=true","fetch_graph":"https://pith.science/api/pith-number/AIPTQ63NI2WFYE33MHJR6KN66O/graph.json","fetch_events":"https://pith.science/api/pith-number/AIPTQ63NI2WFYE33MHJR6KN66O/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/AIPTQ63NI2WFYE33MHJR6KN66O/action/timestamp_anchor","attest_storage":"https://pith.science/pith/AIPTQ63NI2WFYE33MHJR6KN66O/action/storage_attestation","attest_author":"https://pith.science/pith/AIPTQ63NI2WFYE33MHJR6KN66O/action/author_attestation","sign_citation":"https://pith.science/pith/AIPTQ63NI2WFYE33MHJR6KN66O/action/citation_signature","submit_replication":"https://pith.science/pith/AIPTQ63NI2WFYE33MHJR6KN66O/action/replication_record"}},"created_at":"2026-07-05T09:29:35.500678+00:00","updated_at":"2026-07-05T09:29:35.500678+00:00"}