Pith. sign in

Paper Citation Record · LEDGER

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations

As of 12 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2501.07647.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.07647 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:41:46.943342Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b6e5f49-3636-464c-85e6-6c87130acc2b · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.661593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.661593Z digest=sha256:1dd3241a0d2fb2945d6543ac8f7ce442dc6429c1de9fd92aa2d2d9f55381d308

Observation 98b22be7-b64b-49b6-bb56-d74b7c0df6e8 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.940970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.668067Z digest=sha256:03ba65c8d31f192cc33466f11c5ff29895d557c7d7f14b86800f4e9f1484de30

Observation 631277c9-2b1b-4913-9743-51ef281ec728 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.673882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.673882Z digest=sha256:d26ed6df45849cd96c68cf1e125633e7f89dda193e2fc69e6871bf4e22f8e62e

Observation 982eb44b-35c5-4c2f-bacf-68319494733f · outbound

This paper cites Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.924160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.679977Z digest=sha256:b55c572f17a426bb5be7ad1456c5df992c0891736f6e3303d21110fedd6d8d07

Observation 1c833fd9-240e-4cdc-b39b-776bf5163bc5 · outbound

This paper cites Training-free layout control with cross-attention guidance.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Training-free layout control with cross-attention guidance

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.685336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.685336Z digest=sha256:0dc4de42eb724e68465cd09441d0b3af22b3d797bb32fee5a9411474fd47d249

Observation 3db8ca74-d25a-4698-b9b4-01745dd7cbac · outbound

This paper cites Panda-70m: Captioning 70m videos with multiple cross-modality teachers.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Panda-70m: Captioning 70m videos with multiple cross-modality teachers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.894451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.690463Z digest=sha256:7596c8f4f87555030f0b92010cb26a825e29e79219b5b606fd949771716378c4

Observation 62167b68-ea64-405b-95cd-b3d5386b4fcd · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.695844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.695844Z digest=sha256:c8d51bfbacb41940173a097e3d9ab3e202696d6187fe36496bded98741f02945

Observation 491fb305-9892-4bae-a1d2-045effcab649 · outbound

This paper cites TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.702000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.702000Z digest=sha256:6812c798fb96006cf2562f79987a5d76384f78750cc1619e49ca600699ef52e2

Observation aef14e19-ec4e-4cf6-a00e-429c19d20942 · outbound

This paper cites Layoutgpt: Compositional visual plan- ning and generation with large language models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Layoutgpt: Compositional visual plan- ning and generation with large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.865422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.707689Z digest=sha256:a3cfcf625557d164cc4d68e3c5449d3e5aedd785fab4ef4c58ffcbc59a7f83fd

Observation 6f1c1714-89cf-4df0-a3f1-5849036a0d2a · outbound

This paper cites On the content bias in fr ´echet video distance.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations On the content bias in fr ´echet video distance

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.847242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.712470Z digest=sha256:b31a8925aea54f6c47f3fa2af7056de096af50f0092e19adc380bacf02149983

Observation 6e6a0b24-f824-4afb-b223-96d9176a029d · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.717039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.717039Z digest=sha256:7ead31fe9c21d038958f974156c41dd934d8b55a8d7ccc7f54b79fd8830dd95c

Observation f14832ef-e18d-4d32-a077-8c94087197b6 · outbound

This paper cites Denoising dif- fusion probabilistic models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Denoising dif- fusion probabilistic models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.722487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.722487Z digest=sha256:aa29b7938d7779e6dbe87566bd913196544a90af2f79fd7ecf482845a041d732

Observation a509f378-ab9a-4a66-b701-430e417f8341 · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to-video generation via transformers.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Cogvideo: Large-scale pretraining for text-to-video generation via transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.818003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.727435Z digest=sha256:e281d0e89a73552d3f6d47173466d121a1b0185978b03b27be3549e8c341cd89

Observation b512f7f3-1b04-4d41-99d7-6e620314bd92 · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.732481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.732481Z digest=sha256:7b20c1217657457d2b7a994807deeb0fd602245facb25b70f5599ddeaa64f053

Observation c315cf22-2268-4589-8112-96615752f09f · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations A style-based generator architecture for generative adversarial networks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.801856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.737703Z digest=sha256:aa594c7b3b3121aee3c2dc50b56ae1bfba6bc72828e8c5b871fc38f2175c4e2b

Observation 0ea4ed25-7e74-4691-ac5a-8e870a1ef29e · outbound

This paper cites Open-sora-plan, 2024.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Open-sora-plan, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.785045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.743338Z digest=sha256:50f441169a30270d8be64bb93c42b36146193e0d39523161e28db980544f3f0f

Observation f326661a-2d8c-4cec-935a-a19fc11da731 · outbound

This paper cites Dense optical tracking: Connecting the dots.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Dense optical tracking: Connecting the dots

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.768539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.748409Z digest=sha256:ce4ae78a14c9a562b6bfb503066cc24503229540ee8beb263295db14bfa0cdee

Observation 4f48b987-b94a-4382-80d7-4a9ce3047b56 · outbound

This paper cites T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.753046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.753046Z digest=sha256:3840631fd502d8eb3f9b293071d35f59195d04ddbfe6a11e382cabda5ebf8995

Observation 28659284-8aa8-4c06-ab70-20953d6fe2ce · outbound

This paper cites TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.757694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.757694Z digest=sha256:dae04c165005a47cf6255293fe4bdf105b96ee129f564513b416d49d55479eb9

Observation 2205b1eb-5f02-41b9-bafa-78dd3044cfba · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Gligen: Open-set grounded text-to-image generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.751990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.763002Z digest=sha256:0bc85f3128ab94087acb7f2c133b98aedf39fb3ce0574d1684b9b735dd103d33

Observation 63ba12ee-8134-4f6c-821f-c4423f069a4f · outbound

This paper cites Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.735277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.767551Z digest=sha256:c93a6a430ccda5c4dd14df5ac191f66400cb1ce14ae332ebc92ea5b97d55cf71

Observation 98d54c78-f21b-462e-af0b-3d79843d08ec · outbound

This paper cites Llm-grounded video diffusion models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Llm-grounded video diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.718723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.772368Z digest=sha256:4bb4b79f163c476f716a6ecb73500e82e2ddafa7d4d1a1ca34a9d69493ad92c4

Observation 79d97a22-8d0e-4683-a42c-1cbf14c23562 · outbound

This paper cites VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.777075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.777075Z digest=sha256:8200b5f39f46697e06d286a5445aec57afe385718932048d0fb067abf703971b

Observation 1ca446ca-1ebd-4616-90ab-485c5903ad16 · outbound

This paper cites MotionClone: Training-Free Motion Cloning for Controllable Video Generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations MotionClone: Training-Free Motion Cloning for Controllable Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.782104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.782104Z digest=sha256:a8a6f9c7f0aa9aa6b26cc8eeff3f065fc3e7f461c8acdae98c48562fccb1b4a6

Observation 9cd566d9-8960-433c-84e1-b57251091fef · outbound

This paper cites 9 Blobgen-3d: Compositional 3d-consistent freeview image generation with 3d blobs.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations 9 Blobgen-3d: Compositional 3d-consistent freeview image generation with 3d blobs

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.698845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.786747Z digest=sha256:5b3a024668dffd16eb147fa67a8b0b77560c5f843b35ed42d024528d935b37ae

Observation e8486c0f-a363-4208-a11c-460b454bed70 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.679691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.791106Z digest=sha256:0dd19b489caac9413b02beeada829ffe5cbf97733c8766e8dbd8bb5562391edd

Observation 0e5f15a6-bbf0-4b47-9d48-ff02b26a31f0 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.795590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.795590Z digest=sha256:71dbd0ab6f5690cc94f45f65a4b7f11bdab8a17f9034f14d29bac2c96f7a9e5a

Observation 725a678a-e2aa-42c1-bcc8-e348297e1adb · outbound

This paper cites Evalcrafter: Benchmarking and eval- uating large video generation models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Evalcrafter: Benchmarking and eval- uating large video generation models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.662749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.801884Z digest=sha256:600cf742b11ca12c1cf1be811e30961260c4ed88da00a76df3ff92d1b26500ce

Observation 935fbd6d-8b40-454a-9c1a-3a63c044764a · outbound

This paper cites Fetv: A bench- mark for fine-grained evaluation of open-domain text-to- video generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Fetv: A bench- mark for fine-grained evaluation of open-domain text-to- video generation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.646108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.807182Z digest=sha256:7f753030950433672114594d48d0b83f551b5a237ecdf9a1985d657638f8821b

Observation 60011944-a7c1-480e-9db9-e76d8c7bf9bd · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.812519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.812519Z digest=sha256:b4e21e1d004e9b8aac547fcf2db08ac362bc47f12929c915926b983db5205f6b

Observation 05fafea7-0cf5-4277-9f27-4771e99299a0 · outbound

This paper cites Compositional text-to-image gen- eration with dense blob representations.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Compositional text-to-image gen- eration with dense blob representations

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.628927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.817879Z digest=sha256:cef1c0d0081a67bf2ea0c0215b6af375eabc1db1ca7f70a07f751919311d50b0

Observation 2d717236-49e6-4c5e-ba08-857d2cbaa3eb · outbound

This paper cites Scalable diffusion models with transformers.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Scalable diffusion models with transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.822782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.822782Z digest=sha256:85accf3ca6ced9849aa0922f35313d5f44d1285c6f2c26bbac6882f7db2b3cd7

Observation 13761c41-a74a-434e-8e8f-72247be31bd4 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations SAM 2: Segment Anything in Images and Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.827788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.827788Z digest=sha256:60907a2aaa4a5eda1b09d49d10e15fdb3ca6ca49ba7f5bd6c2bb18b099d62e95

Observation 6873bc93-5efe-403c-87a4-62149957d438 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations High-resolution image synthesis with latent diffusion models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.833115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.833115Z digest=sha256:8915c692b0f9f603cd6511f123713502cbd2e89eebe8628146c13edfd918866e

Observation 62089b14-6a46-4765-b1c9-f1f6b602a0d2 · outbound

This paper cites T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.838351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.838351Z digest=sha256:1720c8a2d209b98d81fd1318b901ee44f5d08fa60b2c80f24c336cc98c92c618

Observation 5aa572ee-3d91-417c-9f34-4d0026793aa7 · outbound

This paper cites VidGen-1M: A Large-Scale Dataset for Text-to-video Generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.844463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.844463Z digest=sha256:e52bb2fa24cd42b633457dd4bb3ebb2d5c415d2a61a55902ac4f54c26d1beb6c

Observation 14c488e3-40ff-4233-a5a9-a8794e84021b · outbound

This paper cites Fourier features let networks learn high frequency functions in low dimen- sional domains.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Fourier features let networks learn high frequency functions in low dimen- sional domains

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.850622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.850622Z digest=sha256:2b7d2859e31ecdccfd67d240f817e5e401d94708e61d10b348b73642ee697dcb

Observation ac317e7e-1472-4e97-b502-6692c27daa47 · outbound

This paper cites VideoTetris: Towards Compositional Text-to-Video Generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations VideoTetris: Towards Compositional Text-to-Video Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.855840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.855840Z digest=sha256:15692a33a05767b0c769590ba0810db51642ba2afcbebb3fe6d40950b786184e

Observation 6fa053de-4929-4f99-af7b-6c46cfe7c4d7 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.861690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.861690Z digest=sha256:6ffb2337e39ab7270734599b30b0b5f97e697a2bfa198bebd9ec4950872f37db

Observation 42f5a13b-4a3e-41a6-840e-c7eb99a4043d · outbound

This paper cites Score-based generative modeling in latent space.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Score-based generative modeling in latent space

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.578327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.867718Z digest=sha256:781ddd1d5f1c3965f8a7a010680004a06c6d111e62e3cd63a10a85000ed89403

Observation 8e568363-7a99-4ff4-88c9-4ab15332b295 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations ModelScope Text-to-Video Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.872600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.872600Z digest=sha256:f5dc02d3b7911651c0ee1dd61c3a234034e80b5daff3c44c5ebefa18a8be1a62

Observation c1ce706c-3b10-4e44-bc29-965b62ff289c · outbound

This paper cites Boximator: Gener- ating rich and controllable motions for video synthesis.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Boximator: Gener- ating rich and controllable motions for video synthesis

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.562145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.877786Z digest=sha256:6cc545889f8d9eaf392366a19606b3a91399db024270da3e82c41c281778e167

Observation 237149c0-cfe4-4a60-b919-c6db8fe2eca8 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.883364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.883364Z digest=sha256:bf2773a3680c3190bcc60e294e3dbd3bf0bf8231d915fb347df22c333a7c6e7f

Observation f9b1cfec-798c-4e18-b41d-b42376453dfd · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Msr-vtt: A large video description dataset for bridging video and language

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.888799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.888799Z digest=sha256:14c164f7bb197e78b95f46287cbe62e4711c12c7f41f276154cf43cc96321571

Observation a6d7aa9d-038b-4415-9fa3-c9886b3c3a7f · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.535970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.893683Z digest=sha256:e81313ab946a4ecf7861ea9fb4ed0a83494e7d9080cfa3f2dceff6ace92fcf2e

Observation 29c53d94-d326-43b2-bc10-7c6520361e72 · outbound

This paper cites Ad- vancing high-resolution video-language representation with large-scale video transcriptions.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Ad- vancing high-resolution video-language representation with large-scale video transcriptions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.519590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.898852Z digest=sha256:da644e2be659594c97bf8a59e886aa9c1a82435a25f21a3e2fb7381b98c78946

Observation ff31b245-ad15-466e-ac4c-245d8d1ddc02 · outbound

This paper cites The 3rd large-scale video object segmentation challenge - video in- stance segmentation track, 2021.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations The 3rd large-scale video object segmentation challenge - video in- stance segmentation track, 2021

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.503207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.903448Z digest=sha256:8aae07258c48bb33bc328c1d30bfcc987f51cd9ad9b3ac7b57fbcba12d2524a9

Observation 1c8caf67-ceb5-4d66-87b9-2a5e3a4fed17 · outbound

This paper cites Compositional Video Generation as Flow Equalization.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Compositional Video Generation as Flow Equalization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.907745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.907745Z digest=sha256:b1e3b1dcf5453c360a80a433f2ee34c546344f7fa68a65c2c3efd4f75f78ca7d

Observation b362f654-ab1f-418b-a79a-7247e9727bb1 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.912771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.912771Z digest=sha256:50a4edc01be3c34f4eea5ed9e3d9acc56e0c8fe8903cccb40c9d727bc857f735

Observation d215340d-085f-45be-b3e4-e7d42fc77633 · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d in- door scenes.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Scannet++: A high-fidelity dataset of 3d in- door scenes

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.486813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.917408Z digest=sha256:9e4ca6eba33a6a9ff73b9425b37a59510a0923f8bcb8e6438b30b785344cb999

Observation 6ae457b4-859d-4308-aebf-f18053923296 · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Show-1: Marrying pixel and latent diffusion models for text-to-video generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T20:41:46.921729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:41:46.921729Z digest=sha256:a083d20091c5a3da9df540b1bed63f8a22f490f0bff56bf22e8e59bd46542340

Observation 127e9784-ed54-430e-b1b0-ff3e5fc55929 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Adding conditional control to text-to-image diffusion models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.458369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.926761Z digest=sha256:084f2cbb207e966863d8f1e2878accfe79cf59e4a1952bea26174f083483447d

Observation 51f1e579-ef71-4ae9-9ec9-8cbd61e9839d · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Llava- next: A strong zero-shot video understanding model, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.442037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.931591Z digest=sha256:7628e8dde48dbd205c21f3f4ff9ef0026f2c2bdd4a298e589b79c326ff9b8f16

Observation 5c3d4a12-2050-450f-81be-34389c94427f · outbound

This paper cites both foreground and background.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations both foreground and background

Reference 54

Resolution
verified exact
raw_fallback, observed 2026-08-10T20:41:47.062580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.937177Z digest=sha256:2081b6a2530dddfc29b91cab1eca0afeb20b6acab91d17443b8a313d64703fcb

Observation 87f0d80a-6b97-42d8-b5aa-e73017baeb4f · outbound

This paper cites Frame0”: “Object2.

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations Frame0”: “Object2

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:41:47.424318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:41:46.943342Z digest=sha256:3cea0cddc4ae7cf0f28e617f32c1c00fd6420f530474b6b18435985495d7cba5

Pith citing papers

No inbound Pith citation observations are available.