Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T01:04:35.416676Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2412.18966.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T01:04:35.416676Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4f93188e-2dfe-400a-946b-8444a5f851d3 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d506507-06d0-4fd2-8408-ee526a2c4173 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Lumiere: A space-time dif- fusion model for video generation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 65264e48-4f78-415f-93cf-24f23abd78c0 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Improving image generation with better captions
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7acdce46-e9ff-4954-84f5-bb8420c2d8ac · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Align your latents: High-resolution video synthesis with la- tent diffusion models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1071c5b4-43d3-4666-ab46-58dc859088b2 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement In- structpix2pix: Learning to follow image editing instructions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f99b4b2-d725-4e7b-bdcc-969ce323ba1b · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Video generation models as world simulators
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86ddf32c-274e-4583-bff1-365c83fb2e51 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9afefb62-49e9-4ccb-be21-a32ee5563526 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf650a6b-a6d3-4881-a51a-48b5312e9fff · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image syn- thesis
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5e5f47e3-c7d1-43b0-a32b-271cabe9a253 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Panda-70m: Captioning 70m videos with multiple cross-modality teachers
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1857c1e7-f202-4dc7-b3bc-572ceb3da728 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Contin- ual pre-training mitigates forgetting in language and vision
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d5edb5ab-538e-4b8e-ae38-07867dbf59e8 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement A continual learning survey: Defying for- getting in classification tasks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6115f435-d074-4e1d-a2e9-6d24a305a36d · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Irc- gan: Introspective recurrent convolutional gan for text-to- video generation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8b915457-2b3e-463a-8ebf-507109f939a3 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Catastrophic forgetting in connectionist networks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eb0f6510-8aa8-4672-9a72-3c07a375f1bb · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Videostu- dio: Generating consistent-content and multi-scene videos
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c5a220a6-cc78-4066-bc0d-e24c36358e57 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b0f21a0-c469-4180-8caf-89b54035d244 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Preserve your own correlation: A noise prior for video diffusion models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dbf4ee18-95f0-4d83-9c83-80c374b26fb5 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ae13927-af9e-4072-9a43-94d2677e0788 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Animatediff: Animate your personalized text-to- image diffusion models without specific tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2eb497d-0d51-4551-bac0-1f4449f9b56d · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Photorealistic video generation with diffusion models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 54c4ca97-239e-448c-8ed0-cd1d32be011a · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Latent Video Diffusion Models for High-Fidelity Long Video Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65b639a6-8f91-43ad-8743-471ad106b4f3 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c28b58e-12cf-4e74-b374-64c44e730325 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement LLMs Meet Multimodal Generation and Editing: A Survey
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ef79add-b938-4221-b8c6-13b346135884 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Imagen Video: High Definition Video Generation with Diffusion Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6502a7d2-95b9-479f-87be-87b053ff6552 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Video dif- fusion models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b1d7871-dc3a-42d5-9900-de618ddb4c2c · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6df7b4bc-64e9-4b8c-ad73-677bb985d473 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Parameter-efficient transfer learning for nlp
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfa864c4-37bc-4e12-8618-5eb4d23d8d6f · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Lora: Low- rank adaptation of large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e25e9399-74f3-4834-82c9-e216993040ef · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement VBench: Com- prehensive benchmark suite for video generative models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a122b23e-ac11-46a5-a1ea-d759b41a5827 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Simple and Scalable Strategies to Continually Pre-train Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0065cca-0473-4b6e-826a-66aa850c447e · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Continual pre-training of language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 17de3f09-e731-4c82-8bc1-b9407b894cb5 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Open-sora-plan, 2024
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83037e88-3364-4874-9c81-06457a5e886b · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Llava-next: Tack- ling multi-image, video, and 3d in large multimodal models,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99fc34b5-5ff3-4507-b61a-ceae3cb6d0ce · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Video generation from text
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a42560-6736-442a-9285-dbd368053cf5 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Llm-grounded video diffusion models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f69a88bc-708c-4782-a945-7b7a06fc6934 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Latte: Latent Diffusion Transformer for Video Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6ca6569-8f76-452c-8ac1-19b4e607729b · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Snap video: Scaled spatiotemporal transformers for text-to-video synthesis
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b69fd12d-e3c1-4381-8baf-34053845f54d · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Sync-draw: Automatic video generation using deep recurrent attentive architectures
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5128715b-3fb4-45e2-a9ef-496d73edf303 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Jour- neydb: A benchmark for generative image understanding,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f19c022c-0c68-418c-b541-b8aee8e2c3c4 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Scalable diffusion models with transformers
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a78e60a-17bf-43ac-8009-6d75cfc50cf4 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Learning transferable visual models from natural language supervi- sion
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b467d429-bced-4f93-830d-1607b803caee · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3da6630f-6637-41ae-a223-0035c167fc5e · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19552679-5593-49c2-9c43-f928984dc853 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Emu: Generative Pretraining in Multimodality
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4111d41-5972-45a6-9894-bab94533913a · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Generative multimodal mod- els are in-context learners
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c01c39a3-e640-4fb8-9888-02961b128f2f · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Neural discrete representation learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7c2b8c9-8e60-4fe8-81f9-03bfbb476481 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Magicvideo-v2: Multi-stage high-aesthetic video generation, 2024
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 31901252-b885-434a-9e13-e2a4b7f0fde7 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement TRACE: A Comprehensive Benchmark for Continual Learning in Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c90472d-1505-4b1b-8dbc-f0ea85c79723 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Emu3: Next-Token Prediction is All You Need
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef30cbfb-6a67-4517-b548-ff8ac4c01afd · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d6f23ce-6dc9-4cd2-a31e-5259de3f0449 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement LLaMA pro: Progressive LLaMA with block expansion
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cf330d11-9aa7-4b63-85bd-0c51fcab8702 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Vript: A video is worth thousands of words, 2024
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c755604e-2b24-4701-8e21-fe7597d3bb59 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cc165ec-c1aa-4c8b-8db4-733edfaaf277 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Language model beats diffusion-tokenizer is key to visual generation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8f96385c-1f72-4902-8aac-1f7e543c5e7c · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb8924cb-da26-4933-a0ca-c805c9a3a61e · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Open-sora: Democratizing efficient video production for all, 2024
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e05f910e-1c09-4a7a-834a-311f74a36604 · outbound
ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement MagicVideo: Efficient Video Generation With Latent Diffusion Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.