Pith. sign in

Paper Citation Record · LEDGER

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement

As of 11 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2412.18966.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18966 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T01:04:35.416676Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4f93188e-2dfe-400a-946b-8444a5f851d3 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.891030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.891030Z digest=sha256:0b8e42dd00712cc05b182f472cfe3de66c84c7f7d843a07912e89de677d196a0

Observation 2d506507-06d0-4fd2-8408-ee526a2c4173 · outbound

This paper cites Lumiere: A space-time dif- fusion model for video generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Lumiere: A space-time dif- fusion model for video generation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.424442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:34.899673Z digest=sha256:bcbc0214f4e2b9be102d41f86520dd6312b148f73fa5d397805cef35c5d56018

Observation 65264e48-4f78-415f-93cf-24f23abd78c0 · outbound

This paper cites Improving image generation with better captions.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Improving image generation with better captions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.391062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:34.909158Z digest=sha256:f6a513f95dd892082d29da6ebeeb2ca10451c596c5a22fab42d529de50f3fd17

Observation 7acdce46-e9ff-4954-84f5-bb8420c2d8ac · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.916707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.916707Z digest=sha256:9982ae960ed2f85e26706c15504be089bdef1a10d33bb86139c33c38c1085aec

Observation 1071c5b4-43d3-4666-ab46-58dc859088b2 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement In- structpix2pix: Learning to follow image editing instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.924557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.924557Z digest=sha256:d94cd9756685daaa0798e23a9c91f9162c051bf63afa20fdebdda7c0364124a3

Observation 7f99b4b2-d725-4e7b-bdcc-969ce323ba1b · outbound

This paper cites Video generation models as world simulators.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Video generation models as world simulators

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.934310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.934310Z digest=sha256:275240f3aee1c555bc3ee236ea9165710d73ee6effdedfcabd26bf6ccd9618e5

Observation 86ddf32c-274e-4583-bff1-365c83fb2e51 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.943735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.943735Z digest=sha256:af0a9a616c3cde0c71b788a355cf9b315f3894f5eec49527ba4dbfd62c367329

Observation 9afefb62-49e9-4ccb-be21-a32ee5563526 · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:34.954429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:34.954429Z digest=sha256:e35ee6b80f7ded01e198ee0a2cd5ee58551e02877381477e83819f5f43686b15

Observation cf650a6b-a6d3-4881-a51a-48b5312e9fff · outbound

This paper cites Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image syn- thesis.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image syn- thesis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.226592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:34.961354Z digest=sha256:4c2b010a670f4db43ce4d341910a854fce49e61e33d62e5dbb852702973068b9

Observation 5e5f47e3-c7d1-43b0-a32b-271cabe9a253 · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.198557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:34.968168Z digest=sha256:9afa58ca24acfd0f49657928aa1a4bb6f7c9949b3fecec371113a6f3d3aeabe7

Observation 1857c1e7-f202-4dc7-b3bc-572ceb3da728 · outbound

This paper cites Contin- ual pre-training mitigates forgetting in language and vision.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Contin- ual pre-training mitigates forgetting in language and vision

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.141649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:34.977168Z digest=sha256:b99eb9f7da8d4487b5b5088440fdf1d909b0a6b3bbeec33c26cdacb7e4e2557d

Observation d5edb5ab-538e-4b8e-ae38-07867dbf59e8 · outbound

This paper cites A continual learning survey: Defying for- getting in classification tasks.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement A continual learning survey: Defying for- getting in classification tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.090386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:34.986248Z digest=sha256:84223c58130554ba154569594e191199919d904281b3a4e580f0a18ac34ae334

Observation 6115f435-d074-4e1d-a2e9-6d24a305a36d · outbound

This paper cites Irc- gan: Introspective recurrent convolutional gan for text-to- video generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Irc- gan: Introspective recurrent convolutional gan for text-to- video generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.054938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:34.999548Z digest=sha256:7f94ac3b7d75466d078979c02a724294b3634b17c636446b9180d711df8275ef

Observation 8b915457-2b3e-463a-8ebf-507109f939a3 · outbound

This paper cites Catastrophic forgetting in connectionist networks.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Catastrophic forgetting in connectionist networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:37.020898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.009018Z digest=sha256:049c89032e4f8d76087d4a67895a7215cbf1d5f1ef87b16a07f12f66edf316c7

Observation eb0f6510-8aa8-4672-9a72-3c07a375f1bb · outbound

This paper cites Videostu- dio: Generating consistent-content and multi-scene videos.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Videostu- dio: Generating consistent-content and multi-scene videos

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.979001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.020241Z digest=sha256:3d4ecd762c1399b2b992ce2d6a9172908cb6dfe4da4df253b109ca14e88a3240

Observation c5a220a6-cc78-4066-bc0d-e24c36358e57 · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.030224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.030224Z digest=sha256:46324a5f1bf3a13fa692c79f239fb753ca39ccf4a395c81c739976bbb88fa32e

Observation 9b0f21a0-c469-4180-8caf-89b54035d244 · outbound

This paper cites Preserve your own correlation: A noise prior for video diffusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Preserve your own correlation: A noise prior for video diffusion models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.951409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.038115Z digest=sha256:72664d95091d988a4e5465741f9ce8a2a1377d0085eb7893b5158eb96127d84e

Observation dbf4ee18-95f0-4d83-9c83-80c374b26fb5 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.050124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.050124Z digest=sha256:1bd016be1e08c29605143bae4007a0d934959d249b2b6e68025018978dc5021e

Observation 6ae13927-af9e-4072-9a43-94d2677e0788 · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Animatediff: Animate your personalized text-to- image diffusion models without specific tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.056511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.056511Z digest=sha256:deb33aff0846ce87dabbf4e6c65fdecc40a2ff92ec71f025b09542fb1b108694

Observation e2eb497d-0d51-4551-bac0-1f4449f9b56d · outbound

This paper cites Photorealistic video generation with diffusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Photorealistic video generation with diffusion models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.901018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.068618Z digest=sha256:88b93b47082effec77f98a4ea0d6f219e324f6e325e457dbd0a8eab464e4d6fe

Observation 54c4ca97-239e-448c-8ed0-cd1d32be011a · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.077390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.077390Z digest=sha256:53131b0e568bb8ed1dcf44be65e465769171cf448c155d047d9410acc6f6082d

Observation 65b639a6-8f91-43ad-8743-471ad106b4f3 · outbound

This paper cites Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.082441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.082441Z digest=sha256:f59bd10f108d5dcce416c8cc53f3d5dbf271551ce8aaffa0febe70d949a42f60

Observation 4c28b58e-12cf-4e74-b374-64c44e730325 · outbound

This paper cites LLMs Meet Multimodal Generation and Editing: A Survey.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement LLMs Meet Multimodal Generation and Editing: A Survey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.095299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.095299Z digest=sha256:71a742785189522b071f761c979743a720aa76672581d72718a64d45f7067c6a

Observation 0ef79add-b938-4221-b8c6-13b346135884 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Imagen Video: High Definition Video Generation with Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.105992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.105992Z digest=sha256:b6b88a37575d26dd00dc330c2ffa0dbf562bc46539fee068eefdf9e16aa002cf

Observation 6502a7d2-95b9-479f-87be-87b053ff6552 · outbound

This paper cites Video dif- fusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Video dif- fusion models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.115211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.115211Z digest=sha256:d3e21400fda3556990ca9f9e49893fcb09e1b71859b1c30ecf0931dc83f3198f

Observation 6b1d7871-dc3a-42d5-9900-de618ddb4c2c · outbound

This paper cites DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.123565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.123565Z digest=sha256:7946c59b410146cffef837aea283a16d52fa2fa75d23459245e4b365c2e158cb

Observation 6df7b4bc-64e9-4b8c-ad73-677bb985d473 · outbound

This paper cites Parameter-efficient transfer learning for nlp.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Parameter-efficient transfer learning for nlp

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.129333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.129333Z digest=sha256:03555a0e55a942645b74a954a824b02ca871391f2f5bb3f0329d1bcd4ddc09b0

Observation cfa864c4-37bc-4e12-8618-5eb4d23d8d6f · outbound

This paper cites Lora: Low- rank adaptation of large language models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Lora: Low- rank adaptation of large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.814147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.136837Z digest=sha256:e7015ddab0ea95884757aa74afbaa605b01f22f86fdf2ca92a590d55e6b44ae6

Observation e25e9399-74f3-4834-82c9-e216993040ef · outbound

This paper cites VBench: Com- prehensive benchmark suite for video generative models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement VBench: Com- prehensive benchmark suite for video generative models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.774461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.142817Z digest=sha256:eea48848fd7ddcdd0a9d57df8fa1b80ee10ba158a1c5d34f60e7eb2ab652d935

Observation a122b23e-ac11-46a5-a1ea-d759b41a5827 · outbound

This paper cites Simple and Scalable Strategies to Continually Pre-train Large Language Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Simple and Scalable Strategies to Continually Pre-train Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.150244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.150244Z digest=sha256:41da00b2fd2a1db7c8a76e414863882a4c97cf15f1f60c65f6904d162a59104e

Observation e0065cca-0473-4b6e-826a-66aa850c447e · outbound

This paper cites Continual pre-training of language models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Continual pre-training of language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.738658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.170485Z digest=sha256:1ea3e9879dad7e9ddcb8f832c3ae95dabbe1248bc413983b1aeb710aca5f4975

Observation 17de3f09-e731-4c82-8bc1-b9407b894cb5 · outbound

This paper cites Open-sora-plan, 2024.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Open-sora-plan, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.184878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.184878Z digest=sha256:9461bd3d9676ea92ec7a7d623c7e39a49bf0e47a45d99cf63f333c169d7310de

Observation 83037e88-3364-4874-9c81-06457a5e886b · outbound

This paper cites Llava-next: Tack- ling multi-image, video, and 3d in large multimodal models,.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Llava-next: Tack- ling multi-image, video, and 3d in large multimodal models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.197675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.197675Z digest=sha256:b8f636a7339f112605f9e9597ef6eee1de34c31bc361bcff795c17e8e77d632e

Observation 99fc34b5-5ff3-4507-b61a-ceae3cb6d0ce · outbound

This paper cites Video generation from text.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Video generation from text

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.206474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.206474Z digest=sha256:38488393a2e2f4adfb859779e27cc5f973425d556485ab5dca10b7b874cecb3b

Observation b1a42560-6736-442a-9285-dbd368053cf5 · outbound

This paper cites Llm-grounded video diffusion models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Llm-grounded video diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.630561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.223224Z digest=sha256:25934f899505d335ee1228ab746e3df0e59488410b5d401034f548589951b658

Observation f69a88bc-708c-4782-a945-7b7a06fc6934 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Latte: Latent Diffusion Transformer for Video Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.230155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.230155Z digest=sha256:84c6ab1be87decbc1c299f67d37cb7682215ad0669c35a915f1c2b58ee8d597e

Observation e6ca6569-8f76-452c-8ac1-19b4e607729b · outbound

This paper cites Snap video: Scaled spatiotemporal transformers for text-to-video synthesis.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Snap video: Scaled spatiotemporal transformers for text-to-video synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.240530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.240530Z digest=sha256:5030a380664f8dbc2da6bebcb187c6aedc7c842cee381c7dc49b537f8e6eb1c0

Observation b69fd12d-e3c1-4381-8baf-34053845f54d · outbound

This paper cites Sync-draw: Automatic video generation using deep recurrent attentive architectures.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Sync-draw: Automatic video generation using deep recurrent attentive architectures

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.572032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.256364Z digest=sha256:5792ff7bffb438e2e5b7d39def30717b57bafaea3a944783ac91ee4a3e550e44

Observation 5128715b-3fb4-45e2-a9ef-496d73edf303 · outbound

This paper cites Jour- neydb: A benchmark for generative image understanding,.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Jour- neydb: A benchmark for generative image understanding,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.263136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.263136Z digest=sha256:53ddfa2d2dc0d6065d552e4dac2a6dda46e30c6a902e659715ec6bfe0ee03ac6

Observation f19c022c-0c68-418c-b541-b8aee8e2c3c4 · outbound

This paper cites Scalable diffusion models with transformers.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Scalable diffusion models with transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.270449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.270449Z digest=sha256:c52309936666b91819d1a45080a8904c54a464ff00173b2011a52b06364bd180

Observation 8a78e60a-17bf-43ac-8009-6d75cfc50cf4 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Learning transferable visual models from natural language supervi- sion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.278274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.278274Z digest=sha256:10b56541dae833bbd7f29fce1275122fa59f4d47d35e9e34de28f81ab8d19be5

Observation b467d429-bced-4f93-830d-1607b803caee · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.483662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.282958Z digest=sha256:5bef7d67da3daf8aca7f8f0f9cfb1c1d57617d33c6f898ac426f9c5bf4215601

Observation 3da6630f-6637-41ae-a223-0035c167fc5e · outbound

This paper cites T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.287887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.287887Z digest=sha256:0dd8844ea0568e7d553571574eda5a93f3720a8dad9640f9b0c5625d15fc4969

Observation 19552679-5593-49c2-9c43-f928984dc853 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Emu: Generative Pretraining in Multimodality

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.295977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.295977Z digest=sha256:22ed437b37cb321f0c3cba7cbed67b15bb2fd955a4a251b2a97734c0ca026e8a

Observation e4111d41-5972-45a6-9894-bab94533913a · outbound

This paper cites Generative multimodal mod- els are in-context learners.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Generative multimodal mod- els are in-context learners

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.435782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.304811Z digest=sha256:2e8656606ba4423faabecba018963efb05b5ecb48de968278625cc26eebdf0da

Observation c01c39a3-e640-4fb8-9888-02961b128f2f · outbound

This paper cites Neural discrete representation learning.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Neural discrete representation learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.316897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.316897Z digest=sha256:d61defd65726ff78303aa73856d3c84d9d8e2fa34aceb7af8fd4e0d6e772a9c9

Observation c7c2b8c9-8e60-4fe8-81f9-03bfbb476481 · outbound

This paper cites Magicvideo-v2: Multi-stage high-aesthetic video generation, 2024.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Magicvideo-v2: Multi-stage high-aesthetic video generation, 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.384581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.323886Z digest=sha256:c73ff06f82b9dede11f11e6564d4e2ba9d504add545c27044c36473512a4b1a7

Observation 31901252-b885-434a-9e13-e2a4b7f0fde7 · outbound

This paper cites TRACE: A Comprehensive Benchmark for Continual Learning in Large Language Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement TRACE: A Comprehensive Benchmark for Continual Learning in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.334537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.334537Z digest=sha256:be9d2c58c5869eaf4ef45d6efec84b2d077aa6fdd5641f22999300b172dcc32c

Observation 1c90472d-1505-4b1b-8dbc-f0ea85c79723 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Emu3: Next-Token Prediction is All You Need

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.342459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.342459Z digest=sha256:58dae00e2b5231aec59686d1141179ecffd7d1366632ecb3dbeff1164dad75a9

Observation ef30cbfb-6a67-4517-b548-ff8ac4c01afd · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.349663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.349663Z digest=sha256:aa7c2c4f33b22104915cb5df04f2f5a04ce0e5a8509d8f0383ca8185f247762b

Observation 1d6f23ce-6dc9-4cd2-a31e-5259de3f0449 · outbound

This paper cites LLaMA pro: Progressive LLaMA with block expansion.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement LLaMA pro: Progressive LLaMA with block expansion

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.330890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.358531Z digest=sha256:7d3197c245ea4cdf00e17c41f2d2a121a877f76044deb1b0caab685c529a40ca

Observation cf330d11-9aa7-4b63-85bd-0c51fcab8702 · outbound

This paper cites Vript: A video is worth thousands of words, 2024.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Vript: A video is worth thousands of words, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.288520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.367105Z digest=sha256:77874851e826063a95f81d4dfc193814991631d94a6cb0fc1a621f8f76621c4c

Observation c755604e-2b24-4701-8e21-fe7597d3bb59 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.377149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.377149Z digest=sha256:0b52b8afb67ca5c435180bab6e62ca52db8d920353f6cab42648cef013806482

Observation 5cc165ec-c1aa-4c8b-8db4-733edfaaf277 · outbound

This paper cites Language model beats diffusion-tokenizer is key to visual generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Language model beats diffusion-tokenizer is key to visual generation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.250826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.385041Z digest=sha256:aded402c65675785243bf37cad189447d4aa167032300f4ea149293f3d0b04f6

Observation 8f96385c-1f72-4902-8aac-1f7e543c5e7c · outbound

This paper cites Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.396189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.396189Z digest=sha256:7886e3352ec25b37a1e185d003e10c6d30b2530f4ea9b4f5792f2c5f4aa45075

Observation cb8924cb-da26-4933-a0ca-c805c9a3a61e · outbound

This paper cites Open-sora: Democratizing efficient video production for all, 2024.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement Open-sora: Democratizing efficient video production for all, 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T01:04:36.218328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T01:04:35.406134Z digest=sha256:c512f054769ed45b55e7db0ed2f3d58bc1b9331377a68a64d68dbd3bfbf7da4a

Observation e05f910e-1c09-4a7a-834a-311f74a36604 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:35.416676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:35.416676Z digest=sha256:bcfcb767dc1156b080e944ee01f36cda1908f645c8e43e87acc0acc6240c32fd

Pith citing papers

No inbound Pith citation observations are available.